EDBT 2026 Demo / reviewers in the wild / expert
Hila Gonen
dblp:167/5312
· DBLP profile ↗
22ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins, Stephen Mayhew 0002, Marek Suppa, Hila Gonen, Shachar Mirkin, Vasile Florian Pais, Kaja Dobrovoljc, Voula Giouli, Jun Kevin, Eugene Jang, Eungseo Kim, Jeongyeon Seo, Xenophon Gialis, Yuval Pinter |
LREC | 4 |
| 2026 | Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model BehaviorabstractAbstract We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for intervening on data batches – i.e., “rewriting history” – and then retraining model checkpoints over that data to test hypotheses relating data to behavior. Our intervention recipe’s stages are (1) selecting evaluation items from a benchmark that measures model behavior, (2) matching relevant documents to those items, and (3) modifying those documents before retraining and measuring the effects. We demonstrate the utility of our recipe through case studies on factual knowledge acquisition and gender bias in LMs, using both cooccurrence statistics and information retrieval methods to identify documents that might contribute to model behavior. Our results supplement past observational analyses that link cooccurrence to model behavior, while demonstrating that extant methods for identifying relevant training documents do not fully explain an LM’s abilities and biases. Researchers can follow the recipe to test further hypotheses about how training data affects model behavior. Our code is made publicly available to promote future work.1 Rahul Nadkarni, Yanai Elazar, Hila Gonen, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and ModalitiesabstractThe emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful misuse.The core approach to mitigate these concerns is the detection of harmful queries to the model.Current detection approaches are fallible, and are particularly susceptible to attacks that exploit mismatched generalization of model capabilities (e.g., prompts in lowresource languages or prompts provided in non-text modalities such as image and audio).To tackle this challenge, we propose OMNI-GUARD, an approach for detecting harmful prompts across languages and modalities.Our approach (i) identifies internal representations of an LLM/MLLM that are aligned across languages or modalities and then (ii) uses them to build a language-agnostic or modality-agnostic classifier for detecting harmful prompts.OM-NIGUARD improves harmful prompt classification accuracy by 11.57% over the strongest baseline in a multilingual setting, by 20.44% for image-based prompts, and sets a new SOTA for audio-based prompts.By repurposing embeddings computed during generation, OMNI-GUARD is also very efficient (≈ 120× faster than the next fastest baseline).Code and data are available at https://github.com/ vsahil/OmniGuard. Sahil Verma 0003, Keegan E. Hines, Jeff A. Bilmes, Charlotte Siska, Luke Zettlemoyer, Hila Gonen, Chandan Singh |
EMNLP | 6 |
| 2025 | Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language ModelsabstractHila Gonen, Terra Blevins, Alisa Liu, Luke Zettlemoyer, Noah A. Smith. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hila Gonen, Terra Blevins, Alisa Liu, Luke Zettlemoyer, Noah A. Smith |
NAACL (Long Papers) | 1 |
| 2024 | MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingabstractTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia, Luke Zettlemoyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia, Luke Zettlemoyer |
ACL (1) | 3 |
| 2024 | Voices Unheard: NLP Resources and Models for Yorùbá Regional DialectsabstractOrevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Orevaoghene Ahia, Aremu Anuoluwapo, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, Yulia Tsvetkov |
EMNLP | 4 |
| 2024 | Breaking the Curse of Multilinguality with Cross-lingual Expert Language ModelsabstractTerra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer |
EMNLP | 5 |
| 2024 | Universal NER: A Gold-Standard Multilingual Named Entity Recognition BenchmarkabstractStephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin, Nikola Ljubešić, LJ Miranda, Barbara Plank, Arij Riabi, Yuval Pinter. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Stephen Mayhew 0002, Terra Blevins, Shuheng Liu 0002, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Börje Karlsson 0001, Peiqin Lin, Nikola Ljubesic, Lester James V. Miranda, Barbara Plank, Arij Riabi, Yuval Pinter |
NAACL-HLT | 5 |
| 2024 | BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferabstractAkari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Akari Asai, Sneha Reddy Kudugunta, Xinyan Yu 0001, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi |
NAACL-HLT | 5 |
| 2024 | MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationabstractIn multilingual settings, non-Latin scripts and low-resource languages are usually disadvantaged in terms of language models’ utility, efficiency, and cost. Specifically, previous studies have reported multiple modeling biases that the current tokenization algorithms introduce to non-Latin script languages, the main one being over-segmentation. In this work, we propose MAGNET— multilingual adaptive gradient-based tokenization—to reduce over-segmentation via adaptive gradient-based subword tokenization. MAGNET learns to predict segment boundaries between byte tokens in a sequence via sub-modules within the model, which act as internal boundary predictors (tokenizers). Previous gradient-based tokenization methods aimed for uniform compression across sequences by integrating a single boundary predictor during training and optimizing it end-to-end through stochastic reparameterization alongside the next token prediction objective. However, this approach still results in over-segmentation for non-Latin script languages in multilingual settings. In contrast, MAGNET offers a customizable architecture where byte-level sequences are routed through language-script-specific predictors, each optimized for its respective language script. This modularity enforces equitable segmentation granularity across different language scripts compared to previous methods. Through extensive experiments, we demonstrate that in addition to reducing segmentation disparities, MAGNET also enables faster language modeling and improves downstream utility. Orevaoghene Ahia, Sachin Kumar 0009, Hila Gonen, Valentin Hofmann, Tomasz Limisiewicz, Yulia Tsvetkov, Noah A. Smith |
NeurIPS | 3 |
| 2023 | Prompting Language Models for Linguistic StructureabstractAlthough pretrained language models (PLMs) can be prompted to perform a wide range of language tasks, it remains an open question how much this ability comes from generalizable linguistic understanding versus surface-level lexical patterns.To test this, we present a structured prompting approach for linguistic structured prediction tasks, allowing us to perform zero-and few-shot sequence tagging with autoregressive PLMs.We evaluate this approach on part-of-speech tagging, named entity recognition, and sentence chunking, demonstrating strong few-shot performance in all cases.We also find that while PLMs contain significant prior knowledge of task labels due to task leakage into the pretraining corpus, structured prompting can also retrieve linguistic structure with arbitrary labels.These findings indicate that the in-context learning ability and linguistic knowledge of PLMs generalizes beyond memorization of their training data. Terra Blevins, Hila Gonen, Luke Zettlemoyer |
ACL (1) | 2 |
| 2023 | Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsabstractLanguage models have graduated from being research prototypes to commercialized products offered as web APIs, and recent works have highlighted the multilingual capabilities of these products.The API vendors charge their users based on usage, more specifically on the number of "tokens" processed or generated by the underlying language models.What constitutes a token, however, is training data and model dependent with a large variance in the number of tokens required to convey the same information in different languages.In this work, we analyze the effect of this nonuniformity on the fairness of an API's pricing policy across languages.We conduct a systematic analysis of the cost and utility of OpenAI's language model API on multilingual benchmarks in 22 typologically diverse languages.We show evidence that speakers of a large number of the supported languages are overcharged while obtaining poorer results.These speakers tend to also come from regions where the APIs are less affordable to begin with.Through these analyses, we aim to increase transparency around language model APIs' pricing policies and encourage the vendors to make them more equitable. Orevaoghene Ahia, Sachin Kumar 0009, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, Yulia Tsvetkov |
EMNLP | 3 |
| 2023 | XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsabstractLarge multilingual language models typically rely on a single vocabulary shared across 100+ languages.As these models have increased in parameter count and depth, vocabulary size has remained largely unchanged.This vocabulary bottleneck limits the representational capabilities of multilingual models like XLM-R.In this paper, we introduce a new approach for scaling to very large multilingual vocabularies by de-emphasizing token sharing between languages with little lexical overlap and assigning vocabulary capacity to achieve sufficient coverage for each individual language.Tokenizations using our vocabulary are typically more semantically meaningful and shorter compared to XLM-R.Leveraging this improved vocabulary, we train XLM-V, a multilingual language model with a one million token vocabulary.XLM-V outperforms XLM-R on every task we tested on ranging from natural language inference (XNLI), question answering (MLQA, XQuAD, TyDiQA), to named entity recognition (WikiAnn).XLM-V is particularly effective on low-resource language tasks and outperforms XLM-R by 11.2% and 5.8% absolute on MasakhaNER and Americas NLI, respectively. Davis Liang, Hila Gonen, Yuning Mao, Naman Goyal 0001, Marjan Ghazvininejad, Luke Zettlemoyer, Madian Khabsa |
EMNLP | 2 |
| 2022 | Analyzing the Mono- and Cross-Lingual Pretraining Dynamics of Multilingual Language ModelsabstractThe emergent cross-lingual transfer seen in multilingual pretrained models has sparked significant interest in studying their behavior.However, because these analyses have focused on fully trained multilingual models, little is known about the dynamics of the multilingual pretraining process.We investigate when these models acquire their in-language and crosslingual abilities by probing checkpoints taken from throughout XLM-R pretraining, using a suite of linguistic tasks.Our analysis shows that the model achieves high in-language performance early on, with lower-level linguistic skills acquired before more complex ones.In contrast, the point in pretraining when the model learns to transfer cross-lingually differs across language pairs.Interestingly, we also observe that, across many languages and tasks, the final model layer exhibits significant performance degradation over time, while linguistic knowledge propagates to lower layers of the network.Taken together, these insights highlight the complexity of multilingual pretraining and the resulting varied behavior for different languages over time. Terra Blevins, Hila Gonen, Luke Zettlemoyer |
EMNLP | 2 |
| 2021 | Identifying Helpful Sentences in Product ReviewsabstractIftah Gamzu, Hila Gonen, Gilad Kutiel, Ran Levy, Eugene Agichtein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Iftah Gamzu, Hila Gonen, Gilad Kutiel, Ran Levy 0001, Eugene Agichtein |
NAACL-HLT | 2 |
| 2020 | Simple, Interpretable and Stable Method for Detecting Words with Usage Change across CorporaabstractThe problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science.This is commonly approached by training word embeddings on each corpus, aligning the vector spaces, and looking for words whose cosine distance in the aligned space is large.However, these methods often require extensive filtering of the vocabulary to perform well, and-as we show in this work-result in unstable, and hence less reliable, results.We propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word.The method is simple, interpretable and stable.We demonstrate its effectiveness in 9 different setups, considering different corpus splitting criteria (age, gender and profession of tweet authors, time of tweet) and different languages (English, French and Hebrew). Hila Gonen, Ganesh Jawahar, Djamé Seddah, Yoav Goldberg |
ACL | 1 |
| 2020 | Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionabstractThe ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models.We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations.Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space.By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it.While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification. Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, Yoav Goldberg |
ACL | 3 |
| 2020 | Pick a Fight or Bite your Tongue: Investigation of Gender Differences in Idiomatic Language UsageabstractA large body of research on gender-linked language has established foundations regarding crossgender differences in lexical, emotional, and topical preferences, along with their sociological underpinnings.We compile a novel, large and diverse corpus of spontaneous linguistic productions annotated with speakers' gender, and perform a first large-scale empirical study of distinctions in the usage of figurative language between male and female authors.Our analyses suggest that (1) idiomatic choices reflect gender-specific lexical and semantic preferences in general language, (2) men's and women's idiomatic usages express higher emotion than their literal language, with detectable, albeit more subtle, differences between male and female authors along the dimension of dominance compared to similar distinctions in their literal utterances, and (3) contextual analysis of idiomatic expressions reveals considerable differences, reflecting subtle divergences in usage environments, shaped by cross-gender communication styles and semantic biases. Ella Rabinovich, Hila Gonen, Suzanne Stevenson |
COLING | 2 |
| 2019 | How Does Grammatical Gender Affect Noun Representations in Gender-Marking Languages?abstractMany natural languages assign grammatical gender also to inanimate nouns in the language.In such languages, words that relate to the gender-marked nouns are inflected to agree with the noun's gender.We show that this affects the word representations of inanimate nouns, resulting in nouns with the same gender being closer to each other than nouns with different gender.While "embedding debiasing" methods fail to remove the effect, we demonstrate that a careful application of methods that neutralize grammatical gender signals from the words' context when training word embeddings is effective in removing it.Fixing the grammatical gender bias yields a positive effect on the quality of the resulting word embeddings, both in monolingual and crosslingual settings.We note that successfully removing gender signals, while achievable, is not trivial to do and that a language-specific morphological analyzer, together with careful usage of it, are essential for achieving good results. Hila Gonen, Yova Kementchedjhieva, Yoav Goldberg |
CoNLL | 1 |
| 2019 | Language Modeling for Code-Switching: Evaluation, Integration of Monolingual Data, and Discriminative TrainingabstractHila Gonen, Yoav Goldberg. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hila Gonen, Yoav Goldberg |
EMNLP/IJCNLP (1) | 1 |
| 2019 | It's All in the Name: Mitigating Gender Bias with Name-Based Counterfactual Data SubstitutionabstractRowan Hall Maudslay, Hila Gonen, Ryan Cotterell, Simone Teufel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Rowan Hall Maudslay, Hila Gonen, Ryan Cotterell, Simone Teufel |
EMNLP/IJCNLP (1) | 2 |
| 2016 | Semi Supervised Preposition-Sense Disambiguation using Multilingual DataabstractPrepositions are very common and very ambiguous, and understanding their sense is critical for understanding the meaning of the sentence. Supervised corpora for the preposition-sense disambiguation task are small, suggesting a semi-supervised approach to the task. We show that signals from unannotated multilingual data can be used to improve supervised preposition-sense disambiguation. Our approach pre-trains an LSTM encoder for predicting the translation of a preposition, and then incorporates the pre-trained encoder as a component in a supervised classification system, and fine-tunes it for the task. The multilingual signals consistently improve results on two preposition-sense datasets. Hila Gonen, Yoav Goldberg |
COLING | 1 |