Enora Rice

dblp:372/1671 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 77% Language models and text generation · 14% Machine translation · 10%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%
Human-computer interaction and pervasive computing
1 paper
Design research and methods · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological segmentation
1.622025
Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025
TAMS: Translation-Assisted Morphological Segmentation · ACL (1) 2024
Computational social science and digital humanities
language documentation
1.122025
Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text · EMNLP 2024
Natural language and speech › Information extraction and text analysis
morphological analysis
1.012026
Massively Multilingual Joint Segmentation and Glossing · ACL (1) 2026
Natural language and speech › Information extraction and text analysis
computational morphology
0.912025
Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model
0.812024
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text · EMNLP 2024
Natural language and speech › Machine translation
low-resource machine translation
0.312026
Massively Multilingual Joint Segmentation and Glossing · ACL (1) 2026
Design research and methods
user-centered design
0.312025
Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

user study · 2.6fine-tuning · 1.5multilingual modeling · 1.0pre-trained language model · 0.8multilingual pretraining · 0.8multilingual pre-training · 0.8character-level sequence-to-sequence · 0.8
YearPublicationVenuePosition
2026 Massively Multilingual Joint Segmentation and Glossing
abstract
Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria R. Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer
ACL (1)3
2025 From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation
abstract
Many of the world’s languages have insufficient data to train high-performing general neural machine translation (NMT) models, let alone domain-specific models, and often the only available parallel data are small amounts of religious texts. Hence, domain adaptation (DA) is a crucial issue faced by contemporary NMT and has, so far, been underexplored for low-resource languages. In this paper, we evaluate a set of methods from both low-resource NMT and DA in a realistic setting, in which we aim to translate between a high-resource and a low-resource language with access to only: a) parallel Bible data, b) a bilingual dictionary, and c) a monolingual target-domain corpus in the high-resource language. Our results show that the effectiveness of the tested methods varies, with the simplest one, DALI, being most effective. We follow up with a small human evaluation of DALI, which shows that there is still a need for more careful investigation of how to accomplish DA for low-resource NMT.
Ali Marashian, Enora Rice, Luke Gessler, Alexis Palmer, Katharina von der Wense
COLING2
2025 Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation
abstract
Computational morphology has the potential to support language documentation through tasks like morphological segmentation and the generation of Interlinear Glossed Text (IGT).However, our research outputs have seen limited use in real-world language documentation settings.This position paper situates the disconnect between computational morphology and language documentation within a broader misalignment between research and practice in NLP and argues that the field risks becoming decontextualized and ineffectual without systematic integration of User-Centered Design (UCD).To demonstrate how principles from UCD can reshape the research agenda, we present a case study of GlossLM, a stateof-the-art multilingual IGT generation model.Through a small-scale user study with three documentary linguists, we find that, despite strong metric-based performance, the system fails to meet core usability needs in real documentation contexts.These insights raise new research questions around model constraints, label standardization, segmentation, and personalization.We argue that centering users not only produces more effective tools, but surfaces richer, more relevant research directions.
Enora Rice, Katharina von der Wense, Alexis Palmer
EMNLP1
2024 TAMS: Translation-Assisted Morphological Segmentation
abstract
Canonical morphological segmentation is the process of analyzing words into the standard (aka underlying) forms of their constituent morphemes.This is a core task in endangered language documentation, and NLP systems have the potential to dramatically speed up this process.In typical language documentation settings, training data for canonical morpheme segmentation is scarce, making it difficult to train high quality models.However, translation data is often much more abundant, and, in this work, we present a method that attempts to leverage translation data in the canonical segmentation task.We propose a character-level sequence-to-sequence model that incorporates representations of translations obtained from pretrained high-resource monolingual language models as an additional signal.Our model outperforms the baseline in a super-low resource setting but yields mixed results on training splits with more data.Additionally, we find that we can achieve strong performance even without needing difficult-to-obtain word level alignments.While further work is needed to make translations useful in higher-resource settings, our model shows promise in severely resource-constrained settings.
Enora Rice, Ali Marashian, Luke Gessler, Alexis Palmer, Katharina von der Wense
ACL (1)1
2024 GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text
abstract
Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format.However, there are few existing resources providing large amounts of standardized, easily accessible IGT data, limiting their applicability to linguistic research, and making it difficult to use such data in NLP modeling.We compile the largest existing corpus of IGT data from a variety of sources, covering over 450k examples across 1.8k languages, to enable research on crosslingual transfer and IGT generation.We normalize much of our data to follow a standard set of labels across languages.Furthermore, we explore the task of automatically generating IGT in order to aid documentation projects.As many languages lack sufficient monolingual data, we pretrain a large multilingual model on our corpus.We demonstrate the utility of this model by finetuning it on monolingual corpora, outperforming SOTA models by up to 6.6%.Our pretrained model and dataset are available on Hugging Face.
Michael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice, Graham Neubig, Alexis Palmer, Lori S. Levin
EMNLP4