VLDB 2026 Research / reviewers in the wild / expert
Ekaterina Vylomova
dblp:153/5568 · also Katerina Vylomova
· DBLP profile ↗
13ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-4058-5459ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Machine translation · 46% Information extraction and text analysis · 21% Knowledge representation and reasoning · 19% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
domain adaptation for machine translation |
1.0 | 1 | 2026 | Retrieval for User-Centered Translation: Lessons from RAG-based Tools for Low-Resource Domains · SIGIR 2026 |
Natural language and speech › Machine translation
low-resource machine translation |
1.0 | 1 | 2026 | Retrieval for User-Centered Translation: Lessons from RAG-based Tools for Low-Resource Domains · SIGIR 2026 |
Natural language and speech › Machine translation › neural machine translation
retrieval-augmented machine translation |
1.0 | 1 | 2026 | Retrieval for User-Centered Translation: Lessons from RAG-based Tools for Low-Resource Domains · SIGIR 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
spatial language understanding |
1.0 | 1 | 2026 | Where the Cat Sat: A Multilingual Framework for Spatial Language Understanding · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection |
0.8 | 1 | 2024 | A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications · ACL (1) 2024 |
Natural language and speech › Language models and text generation › natural language understanding
multilingual language understanding |
0.3 | 1 | 2026 | Where the Cat Sat: A Multilingual Framework for Spatial Language Understanding · ACL (1) 2026 |
Information retrieval
retrieval-augmented generation |
0.3 | 1 | 2026 | Retrieval for User-Centered Translation: Lessons from RAG-based Tools for Low-Resource Domains · SIGIR 2026 |
Natural language and speech › Language models and text generation › text generation › surface realization
morphological generation |
0.3 | 1 | 2017 | Paradigm Completion for Derivational Morphology · EMNLP 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › semantic relations
lexical relation learning |
0.2 | 1 | 2016 | Take and Took, Gaggle and Goose, Book and Read: Evaluating the Utility of Vector Differences for Lexical Relation Learning · ACL (1) 2016 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 1 | 2016 | Take and Took, Gaggle and Goose, Book and Read: Evaluating the Utility of Vector Differences for Lexical Relation Learning · ACL (1) 2016 |
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation |
0.2 | 1 | 2014 | Classifying Idiomatic and Literal Expressions Using Topic Models and Intensity of Emotions · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 1 | 2014 | Classifying Idiomatic and Literal Expressions Using Topic Models and Intensity of Emotions · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
emotion recognition |
0.1 | 1 | 2014 | Classifying Idiomatic and Literal Expressions Using Topic Models and Intensity of Emotions · EMNLP 2014 |
Methods — techniques the papers use, named apart from their topics
neural machine translation · 2.0large language model post-editing · 2.0in-context learning · 2.0sentiment analysis · 1.5embedding-based semantic change detection · 1.5collocate analysis · 1.5multilingual evaluation · 1.0neural sequence-to-sequence models · 0.3supervised learning · 0.2spectral clustering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where the Cat Sat: A Multilingual Framework for Spatial Language UnderstandingabstractDemian Inostroza, Ekaterina Vylomova, Charles Kemp, Mae Carroll, Wanchun Li, Meladel Mistica. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Demian Inostroza Améstica, Ekaterina Vylomova, Charles Kemp, Mae Carroll, Wanchun Li, Meladel Mistica |
ACL (1) | 2 |
| 2026 | CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan, Rico Sennrich, Eduard H. Hovy, Ekaterina Vylomova |
LREC | 6 |
| 2026 | Retrieval for User-Centered Translation: Lessons from RAG-based Tools for Low-Resource DomainsabstractMachine translation for low-resource languages suffers from domain-imbalanced corpora, causing quality degradation on technical text. However, in-context learning opens the possibility to rely on limited in-domain corpora to inform translation. We present lessons learned from Tulun, a retrieval-augmented system combining neural MT with LLM post-editing, guided by user-configurable translation memories and glossaries. Deployed for medical translation in Timor-Leste (Tetun) and disaster relief translation in Vanuatu (Bislama), the system achieves accuracy improvements over baseline MT by 16.90-22.41 ChrF++ points, while offering rapid adaptability and transparency to end-users. Key recommendations include: domain granularity matters more than broad categories; translation target audience should inform retrieval; and RAG-augmented MT is most effective for languages that lack domain corpora but remain within LLM pretraining distributions. Raphaël Merx, Ekaterina Vylomova |
SIGIR | 2 |
| 2025 | Usage frequency predicts lexicalization across languages
Temuulen Khishigsuren, Francis Mollica, Ekaterina Vylomova, Charles Kemp |
CogSci | 3 |
| 2024 | A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science ApplicationsabstractHistorical linguists have identified multiple forms of lexical semantic change.We present a three-dimensional framework for integrating these forms and a unified computational methodology for evaluating them concurrently.The dimensions represent increases or decreases in semantic 1) sentiment (valence of a target word's collocates), 2) breadth (diversity of contexts in which the target word appears), and 3) intensity (emotional arousal of collocates or the frequency of intensifiers).These dimensions can be complemented by the evaluation of shifts in the frequency of the target words and the thematic content of its collocates.This framework enables lexical semantic change to be mapped economically and systematically and has applications in computational social science.We present an illustrative analysis of semantic shifts in mental health and mental illness in two corpora, demonstrating patterns of semantic change that illuminate contemporary concerns about pathologization, stigma, and concept creep. Naomi Baes, Nick Haslam, Ekaterina Vylomova |
ACL (1) | 3 |
| 2024 | Predicting Human Translation Difficulty with Neural Machine TranslationabstractAbstract Human translators linger on some words and phrases more than others, and predicting this variation is a step towards explaining the underlying cognitive processes. Using data from the CRITT Translation Process Research Database, we evaluate the extent to which surprisal and attentional features derived from a Neural Machine Translation (NMT) model account for reading and production times of human translators. We find that surprisal and attention are complementary predictors of translation difficulty, and that surprisal derived from a NMT model is the single most successful predictor of production duration. Our analyses draw on data from hundreds of translators operating across 13 language pairs, and represent the most comprehensive investigation of human translation difficulty to date. Zheng Wei Lim, Ekaterina Vylomova, Charles Kemp, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 95 |
| 2020 | UniMorph 3.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological paradigms for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. We have implemented several improvements to the extraction pipeline which creates most of our data, so that it is both more complete and more correct. We have added 66 new languages, as well as new parts of speech for 12 languages. We have also amended the schema in several ways. Finally, we present three new community tools: two to validate data for resource creators, and one to make morphological data available from the command line. UniMorph is based at the Center for Language and Speech Processing (CLSP) at Johns Hopkins University in Baltimore, Maryland. This paper details advances made to the schema, tooling, and dissemination of project resources since the UniMorph 2.0 release described at LREC 2018. Arya McCarthy, Christo Kirov, Matteo Grella, Amrit Nidhi, Patrick Xia 0002, Kyle Gorman, Ekaterina Vylomova, Sabrina J. Mielke, Garrett Nicolai, Miikka Silfverberg, Timofey Arkhangelskiy, Nataly Krizhanovsky, Andrew Krizhanovsky, Elena Klyachko, Alexey Sorokin, John Mansfield, Valts Ernstreits, Yuval Pinter, Cassandra L. Jacobs, Ryan Cotterell, Mans Hulden, David Yarowsky |
LREC | 7 |
| 2019 | Weird Inflects but OK: Making Sense of Morphological Generation ErrorsabstractWe conduct a manual error analysis of the CoNLL-SIGMORPHON 2017 Shared Task on Morphological Reinflection.In this task, systems are given a word in citation form (e.g., hug) and asked to produce the corresponding inflected form (e.g., the simple past hugged).This design lets us analyze errors much like we might analyze children's production errors.We propose an error taxonomy and use it to annotate errors made by the top two systems across twelve languages.Many of the observed errors are related to inflectional patterns sensitive to inherent linguistic properties such as animacy or affect; many others are failures to predict truly unpredictable inflectional behaviors.We also find nearly one quarter of the residual "errors" reflect errors in the gold data. Kyle Gorman, Arya McCarthy, Ryan Cotterell, Ekaterina Vylomova, Miikka Silfverberg, Magdalena Markowska |
CoNLL | 4 |
| 2018 | UniMorph 2.0: Universal Morphology
Christo Kirov, Ryan Cotterell, John Sylak-Glassman, Géraldine Walther, Ekaterina Vylomova, Patrick Xia 0002, Manaal Faruqui, Sabrina J. Mielke, Arya McCarthy, Sandra Kübler, David Yarowsky, Jason Eisner, Mans Hulden |
LREC | 5 |
| 2017 | Paradigm Completion for Derivational MorphologyabstractThe generation of complex derived word forms has been an overlooked problem in NLP; we fill this gap by applying neural sequence-to-sequence models to the task.We overview the theoretical motivation for a paradigmatic treatment of derivational morphology, and introduce the task of derivational paradigm completion as a parallel to inflectional paradigm completion.State-of-the-art neural models, adapted from the inflection task, are able to learn a range of derivation patterns, and outperform a non-neural baseline by 16.4%.However, due to semantic, historical, and lexical considerations involved in derivational morphology, future work will be needed to achieve performance parity with inflection-generating systems. Ryan Cotterell, Ekaterina Vylomova, Huda Khayrallah, Christo Kirov, David Yarowsky |
EMNLP | 2 |
| 2016 | Take and Took, Gaggle and Goose, Book and Read: Evaluating the Utility of Vector Differences for Lexical Relation LearningabstractRecent work has shown that simple vector subtraction over word embeddings is surprisingly effective at capturing different lexical relations, despite lacking explicit supervision.Prior work has evaluated this intriguing result using a word analogy prediction formulation and hand-selected relations, but the generality of the finding over a broader range of lexical relation types and different learning settings has not been evaluated.In this paper, we carry out such an evaluation in two learning settings:(1) spectral clustering to induce word relations, and ( 2) supervised learning to classify vector differences into relation types.We find that word embeddings capture a surprising amount of information, and that, under suitable supervised training, vector subtraction generalises well to a broad range of relations, including over unseen lexical items. Ekaterina Vylomova, Laura Rimell, Trevor Cohn, Timothy Baldwin |
ACL (1) | 1 |
| 2014 | Classifying Idiomatic and Literal Expressions Using Topic Models and Intensity of EmotionsabstractWe describe an algorithm for automatic classification of idiomatic and literal expressions.Our starting point is that words in a given text segment, such as a paragraph, that are highranking representatives of a common topic of discussion are less likely to be a part of an idiomatic expression.Our additional hypothesis is that contexts in which idioms occur, typically, are more affective and therefore, we incorporate a simple analysis of the intensity of the emotions expressed by the contexts.We investigate the bag of words topic representation of one to three paragraphs containing an expression that should be classified as idiomatic or literal (a target phrase).We extract topics from paragraphs containing idioms and from paragraphs containing literals using an unsupervised clustering method, Latent Dirichlet Allocation (LDA) (Blei et al., 2003).Since idiomatic expressions exhibit the property of non-compositionality, we assume that they usually present different semantics than the words used in the local topic.We treat idioms as semantic outliers, and the identification of a semantic shift as outlier detection.Thus, this topic representation allows us to differentiate idioms from literals using local semantic contexts.Our results are encouraging. Jing Peng 0001, Anna Feldman, Ekaterina Vylomova |
EMNLP | 3 |