VLDB 2026 Research / reviewers in the wild / expert
Ekaterina Kochmar
dblp:140/3465
· DBLP profile ↗
25ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-3328-1374ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Reward Modeling for AI Tutors in Math Mistake Remediation
Kseniia Petukhova, Ekaterina Kochmar |
LREC | 2 |
| 2025 | KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of KazakhstanabstractMukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov, Maiya Goloburda, Akhmed Sakip, Zhuohan Xie, Yuxia Wang, Bekassyl Syzdykov, Nurkhan Laiyk, Alham Fikri Aji, Ekaterina Kochmar, Preslav Nakov, Fajri Koto. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov, Maiya Goloburda, Akhmed Sakip, Zhuohan Xie, Yuxia Wang 0003, Bekassyl Syzdykov, Nurkhan Laiyk, Alham Fikri Aji, Ekaterina Kochmar, Preslav Nakov, Fajri Koto |
ACL (1) | 12 |
| 2025 | What Makes Cryptic Crosswords Challenging for LLMs?abstractCryptic crosswords are puzzles that rely on general knowledge and the solver’s ability to manipulate language on different levels, dealing with various types of wordplay. Previous research suggests that solving such puzzles is challenging even for modern NLP models, including Large Language Models (LLMs). However, there is little to no research on the reasons for their poor performance on this task. In this paper, we establish the benchmark results for three popular LLMs: Gemma2, LLaMA3 and ChatGPT, showing that their performance on this task is still significantly below that of humans. We also investigate why these models struggle to achieve superior performance. We release our code and introduced datasets at https://github.com/bodasadallah/decrypting-crosswords. Abdelrahman Boda Sadallah, Daria Kotova, Ekaterina Kochmar |
COLING | 3 |
| 2025 | UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency AssessmentabstractJoseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds 0001, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi |
EMNLP | 11 |
| 2025 | LLMs cannot spot math errors, even when allowed to peek into the solutionabstractLarge language models (LLMs) demonstrate remarkable performance on math word problems, yet they have been shown to struggle with meta-reasoning tasks such as identifying errors in student solutions.In this work, we investigate the challenge of locating the first error step in stepwise solutions using two error reasoning datasets: VtG and PRM800K.Our experiments show that state-of-the-art LLMs struggle to locate the first error step in student solutions even when given access to the reference solution.To that end, we propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student's solution, which helps improve performance. KV Aditya Srivatsa, Kaushal Kumar Maurya, Ekaterina Kochmar |
EMNLP | 3 |
| 2025 | Generative AI for Early Grade Story Generation Using a Self-Reflective Approach
Taufiq Syed, Aadhith Shankarnarayanan, Yara Kaddoura, Salsabeel Y. Shapsough, Imran A. Zualkernan, Ekaterina Kochmar |
ICALT | 6 |
| 2025 | Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI TutorsabstractKaushal Kumar Maurya, Kv Aditya Srivatsa, Kseniia Petukhova, Ekaterina Kochmar. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kaushal Kumar Maurya, KV Aditya Srivatsa, Kseniia Petukhova, Ekaterina Kochmar |
NAACL (Long Papers) | 4 |
| 2024 | How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational QuizzesabstractQuestion generation (QG) is a natural language processing task with an abundance of potential benefits and use cases in the educational domain. In order for this potential to be realized, QG systems must be designed and validated with pedagogical needs in mind. However, little research has assessed or designed QG approaches with the input of real teachers or students. This paper applies a large language model-based QG approach where questions are generated with learning goals derived from Bloom's taxonomy. The automatically generated questions are used in multiple experiments designed to assess how teachers use them in practice. The results demonstrate that teachers prefer to write quizzes with automatically generated questions, and that such quizzes have no loss in quality compared to handwritten versions. Further, several metrics indicate that automatically generated questions can even improve the quality of the quizzes created, showing the promise for large scale use of QG in the classroom setting. Sabina Elkins, Ekaterina Kochmar, Jackie Chi Kit Cheung, Iulian Serban |
AAAI | 2 |
| 2024 | REFeREE: A REference-FREE Model-Based Metric for Text SimplificationabstractText simplification lacks a universal standard of quality, and annotated reference simplifications are scarce and costly. We propose to alleviate such limitations by introducing REFeREE, a reference-free model-based metric with a 3-stage curriculum. REFeREE leverages an arbitrarily scalable pretraining stage and can be applied to any quality standard as long as a small number of human annotations are available. Our experiments show that our metric outperforms existing reference-based metrics in predicting overall ratings and reaches competitive and consistent performance in predicting specific ratings while requiring no reference simplifications at inference time. Ekaterina Kochmar |
LREC/COLING | 2 |
| 2023 | BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine LanguagesabstractCurrent research on automatic readability assessment (ARA) has focused on improving the performance of models in high-resource languages such as English.In this work, we introduce and release BASAHACORPUS as part of an initiative aimed at expanding available corpora and baseline models for readability assessment in lower resource languages in the Philippines.We compiled a corpus of short fictional narratives written in Hiligaynon, Minasbate, Karaya, and Rinconada-languages belonging to the Central Philippine family tree subgroup-to train ARA models using surface-level, syllablepattern, and n-gram overlap features.We also propose a new hierarchical cross-lingual modeling approach that takes advantage of a language's placement in the family tree to increase the amount of available training data.Our study yields encouraging results that support previous work showcasing the efficacy of cross-lingual models in low-resource settings, as well as similarities in highly informative linguistic features for mutually intelligible languages.1 Joseph Marvin Imperial, Ekaterina Kochmar |
EMNLP | 2 |
| 2022 | Raising Student Completion Rates with Adaptive Curriculum and Contextual Bandits
Robert Belfer, Ekaterina Kochmar, Iulian Serban |
AIED (1) | 2 |
| 2021 | Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor SystemsabstractWe explore creating automated, personalized feedback in an intelligent tutoring system (ITS). Our goal is to pinpoint correct and incorrect concepts in student answers in order to achieve better student learning gains. Although automatic methods for providing personalized feedback exist, they do not explicitly inform students about which concepts in their answers are correct or incorrect. Our approach involves decomposing students answers using neural discourse segmentation and classification techniques. This decomposition yields a relational graph over all discourse units covered by the reference solutions and student answers. We use this inferred relational graph structure and a neural classifier to match student answers with reference solutions and generate personalized feedback. Although the process is completely automated and data-driven, the personalized feedback generated is highly contextual, domain-aware and effectively targets each student's misconceptions and knowledge gaps. We test our method in a dialogue-based ITS and demonstrate that our approach results in high-quality feedback and significantly improved student learning gains. Matt Grenander, Robert Belfer, Ekaterina Kochmar, Iulian Serban, François St-Hilaire, Jackie Chi Kit Cheung |
AAAI | 3 |
| 2021 | A Comparative Study of Learning Outcomes for Online Learning Platforms
François St-Hilaire, Nathan Burns, Robert Belfer, Muhammad Shayan, Ariella Smofsky, Dung Do Vu, Antoine Frau, Joseph Potochny, Farid Faraji, Vincent Pavero, Neroli Ko, Ansona Onyi Ching, Sabina Elkins, Anush Stepanyan, Adela Matajova, Laurent Charlin, Yoshua Bengio, Iulian Serban, Ekaterina Kochmar |
AIED (2) | 19 |
| 2021 | A New Readability Assessment Tool
Rebecca Watson, Ekaterina Kochmar |
EDM | 2 |
| 2021 | Word Complexity is in the Eye of the BeholderabstractSian Gooding, Ekaterina Kochmar, Seid Muhie Yimam, Chris Biemann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Sian Gooding, Ekaterina Kochmar, Seid Muhie Yimam, Chris Biemann |
NAACL-HLT | 2 |
| 2020 | Automated Personalized Feedback Improves Learning Gains in An Intelligent Tutoring System
Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Iulian Serban, Joelle Pineau |
AIED (2) | 1 |
| 2020 | A Large-Scale, Open-Domain, Mixed-Interface Dialogue-Based ITS for STEM
Iulian Serban, Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Joelle Pineau, Aaron C. Courville, Laurent Charlin, Yoshua Bengio |
AIED (2) | 3 |
| 2020 | Detecting Multiword Expression Type Helps Lexical Complexity AssessmentabstractMultiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature. Multiple NLP applications have been shown to benefit from MWE identification, however the research on lexical complexity of MWEs is still an under-explored area. In this work, we re-annotate the Complex Word Identification Shared Task 2018 dataset of Yimam et al. (2017), which provides complexity scores for a range of lexemes, with the types of MWEs. We release the MWE-annotated dataset with this paper, and we believe this dataset represents a valuable resource for the text simplification community. In addition, we investigate which types of expressions are most problematic for native and non-native readers. Finally, we show that a lexical complexity assessment system benefits from the information about MWE types. Ekaterina Kochmar, Sian Gooding, Matthew Shardlow |
LREC | 1 |
| 2020 | SeCoDa: Sense Complexity DatasetabstractThe Sense Complexity Dataset (SeCoDa) provides a corpus that is annotated jointly for complexity and word senses. It thus provides a valuable resource for both word sense disambiguation and the task of complex word identification. The intention is that this dataset will be used to identify complexity at the level of word senses rather than word tokens. For word sense annotation SeCoDa uses a hierarchical scheme that is based on information available in the Cambridge Advanced Learner’s Dictionary. This way we can offer more coarse-grained senses than directly available in WordNet. David Strohmaier, Sian Gooding, Shiva Taslimipoor, Ekaterina Kochmar |
LREC | 4 |
| 2019 | Complex Word Identification as a Sequence Labelling TaskabstractComplex Word Identification (CWI) is concerned with detection of words in need of simplification and is a crucial first step in a simplification pipeline.It has been shown that reliable CWI systems considerably improve text simplification.However, most CWI systems to date address the task on a word-by-word basis, not taking the context into account.In this paper, we present a novel approach to CWI based on sequence modelling.Our system is capable of performing CWI in context, does not require extensive feature engineering and outperforms state-of-the-art systems on this task. Sian Gooding, Ekaterina Kochmar |
ACL (1) | 2 |
| 2019 | Recursive Context-Aware Lexical SimplificationabstractSian Gooding, Ekaterina Kochmar. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sian Gooding, Ekaterina Kochmar |
EMNLP/IJCNLP (1) | 2 |
| 2017 | Classification of Twitter Accounts into Automated Agents and Human UsersabstractOnline social networks (OSNs) have seen a remarkable rise in the presence of surreptitious automated accounts. Massive human user-base and business-supportive operating model of social networks (such as Twitter) facilitates the creation of automated agents. In this paper we outline a systematic methodology and train a classifier to categorise Twitter accounts into 'automated' and 'human' users. To improve classification accuracy we employ a set of novel steps. First, we divide the dataset into four popularity bands to compensate for differences in types of accounts. Second, we create a large ground truth dataset using human annotations and extract relevant features from raw tweets. To judge accuracy of the procedure we calculate agreement among human annotators as well as with a bot detection research tool. We then apply a Random Forests classifier that achieves an accuracy close to human agreement. Finally, as a concluding step we perform tests to measure the efficacy of our results. Zafar Gilani, Ekaterina Kochmar, Jon Crowcroft |
ASONAM | 2 |
| 2016 | Cross-Lingual Lexico-Semantic Transfer in Language LearningabstractLexico-semantic knowledge of our native language provides an initial foundation for second language learning.In this paper, we investigate whether and to what extent the lexico-semantic models of the native language (L1) are transferred to the second language (L2).Specifically, we focus on the problem of lexical choice and investigate it in the context of three typologically diverse languages: Russian, Spanish and English.We show that a statistical semantic model learned from L1 data improves automatic error detection in L2 for the speakers of the respective L1.Finally, we investigate whether the semantic model learned from a particular L1 is portable to other, typologically related languages. Ekaterina Kochmar, Ekaterina Shutova |
ACL (1) | 1 |
| 2016 | 'Calling on the classical phone': a distributional model of adjective-noun errors in learners' EnglishabstractIn this paper we discuss three key points related to error detection (ED) in learners’ English. We focus on content word ED as one of the most challenging tasks in this area, illustrating our claims on adjective–noun (AN) combinations. In particular, we (1) investigate the role of context in accurately capturing semantic anomalies and implement a system based on distributional topic coherence, which achieves state-of-the-art accuracy on a standard test set; (2) thoroughly investigate our system’s performance across individual adjective classes, concluding that a class-dependent approach is beneficial to the task; (3) discuss the data size bottleneck in this area, and highlight the challenges of automatic error generation for content words. Aurélie Herbelot, Ekaterina Kochmar |
COLING | 2 |
| 2014 | Detecting Learner Errors in the Choice of Content Words Using Compositional Distributional Semantics
Ekaterina Kochmar, Ted Briscoe |
COLING | 1 |