VLDB 2026 Research / reviewers in the wild / expert
Enora Rice
dblp:372/1671
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Information extraction and text analysis · 77% Language models and text generation · 14% Machine translation · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Design research and methods · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological segmentation |
1.6 | 2 | 2025 | Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025 TAMS: Translation-Assisted Morphological Segmentation · ACL (1) 2024 |
Computational social science and digital humanities
language documentation |
1.1 | 2 | 2025 | Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025 GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
morphological analysis |
1.0 | 1 | 2026 | Massively Multilingual Joint Segmentation and Glossing · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
computational morphology |
0.9 | 1 | 2025 | Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025 |
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model |
0.8 | 1 | 2024 | GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text · EMNLP 2024 |
Natural language and speech › Machine translation
low-resource machine translation |
0.3 | 1 | 2026 | Massively Multilingual Joint Segmentation and Glossing · ACL (1) 2026 |
Design research and methods
user-centered design |
0.3 | 1 | 2025 | Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
user study · 2.6fine-tuning · 1.5multilingual modeling · 1.0pre-trained language model · 0.8multilingual pretraining · 0.8multilingual pre-training · 0.8character-level sequence-to-sequence · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Massively Multilingual Joint Segmentation and GlossingabstractMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria R. Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer |
ACL (1) | 3 |
| 2025 | From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine TranslationabstractMany of the world’s languages have insufficient data to train high-performing general neural machine translation (NMT) models, let alone domain-specific models, and often the only available parallel data are small amounts of religious texts. Hence, domain adaptation (DA) is a crucial issue faced by contemporary NMT and has, so far, been underexplored for low-resource languages. In this paper, we evaluate a set of methods from both low-resource NMT and DA in a realistic setting, in which we aim to translate between a high-resource and a low-resource language with access to only: a) parallel Bible data, b) a bilingual dictionary, and c) a monolingual target-domain corpus in the high-resource language. Our results show that the effectiveness of the tested methods varies, with the simplest one, DALI, being most effective. We follow up with a small human evaluation of DALI, which shows that there is still a need for more careful investigation of how to accomplish DA for low-resource NMT. Ali Marashian, Enora Rice, Luke Gessler, Alexis Palmer, Katharina von der Wense |
COLING | 2 |
| 2025 | Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language DocumentationabstractComputational morphology has the potential to support language documentation through tasks like morphological segmentation and the generation of Interlinear Glossed Text (IGT).However, our research outputs have seen limited use in real-world language documentation settings.This position paper situates the disconnect between computational morphology and language documentation within a broader misalignment between research and practice in NLP and argues that the field risks becoming decontextualized and ineffectual without systematic integration of User-Centered Design (UCD).To demonstrate how principles from UCD can reshape the research agenda, we present a case study of GlossLM, a stateof-the-art multilingual IGT generation model.Through a small-scale user study with three documentary linguists, we find that, despite strong metric-based performance, the system fails to meet core usability needs in real documentation contexts.These insights raise new research questions around model constraints, label standardization, segmentation, and personalization.We argue that centering users not only produces more effective tools, but surfaces richer, more relevant research directions. Enora Rice, Katharina von der Wense, Alexis Palmer |
EMNLP | 1 |
| 2024 | TAMS: Translation-Assisted Morphological SegmentationabstractCanonical morphological segmentation is the process of analyzing words into the standard (aka underlying) forms of their constituent morphemes.This is a core task in endangered language documentation, and NLP systems have the potential to dramatically speed up this process.In typical language documentation settings, training data for canonical morpheme segmentation is scarce, making it difficult to train high quality models.However, translation data is often much more abundant, and, in this work, we present a method that attempts to leverage translation data in the canonical segmentation task.We propose a character-level sequence-to-sequence model that incorporates representations of translations obtained from pretrained high-resource monolingual language models as an additional signal.Our model outperforms the baseline in a super-low resource setting but yields mixed results on training splits with more data.Additionally, we find that we can achieve strong performance even without needing difficult-to-obtain word level alignments.While further work is needed to make translations useful in higher-resource settings, our model shows promise in severely resource-constrained settings. Enora Rice, Ali Marashian, Luke Gessler, Alexis Palmer, Katharina von der Wense |
ACL (1) | 1 |
| 2024 | GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed TextabstractLanguage documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format.However, there are few existing resources providing large amounts of standardized, easily accessible IGT data, limiting their applicability to linguistic research, and making it difficult to use such data in NLP modeling.We compile the largest existing corpus of IGT data from a variety of sources, covering over 450k examples across 1.8k languages, to enable research on crosslingual transfer and IGT generation.We normalize much of our data to follow a standard set of labels across languages.Furthermore, we explore the task of automatically generating IGT in order to aid documentation projects.As many languages lack sufficient monolingual data, we pretrain a large multilingual model on our corpus.We demonstrate the utility of this model by finetuning it on monolingual corpora, outperforming SOTA models by up to 6.6%.Our pretrained model and dataset are available on Hugging Face. Michael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice, Graham Neubig, Alexis Palmer, Lori S. Levin |
EMNLP | 4 |