Ekaterina Kochmar

dblp:140/3465 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-3328-1374ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Towards Reward Modeling for AI Tutors in Math Mistake Remediation
Kseniia Petukhova, Ekaterina Kochmar
LREC2
2025 KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
abstract
Mukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov, Maiya Goloburda, Akhmed Sakip, Zhuohan Xie, Yuxia Wang, Bekassyl Syzdykov, Nurkhan Laiyk, Alham Fikri Aji, Ekaterina Kochmar, Preslav Nakov, Fajri Koto. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Mukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov, Maiya Goloburda, Akhmed Sakip, Zhuohan Xie, Yuxia Wang 0003, Bekassyl Syzdykov, Nurkhan Laiyk, Alham Fikri Aji, Ekaterina Kochmar, Preslav Nakov, Fajri Koto
ACL (1)12
2025 What Makes Cryptic Crosswords Challenging for LLMs?
abstract
Cryptic crosswords are puzzles that rely on general knowledge and the solver’s ability to manipulate language on different levels, dealing with various types of wordplay. Previous research suggests that solving such puzzles is challenging even for modern NLP models, including Large Language Models (LLMs). However, there is little to no research on the reasons for their poor performance on this task. In this paper, we establish the benchmark results for three popular LLMs: Gemma2, LLaMA3 and ChatGPT, showing that their performance on this task is still significantly below that of humans. We also investigate why these models struggle to achieve superior performance. We release our code and introduced datasets at https://github.com/bodasadallah/decrypting-crosswords.
Abdelrahman Boda Sadallah, Daria Kotova, Ekaterina Kochmar
COLING3
2025 UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
abstract
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds 0001, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
EMNLP11
2025 LLMs cannot spot math errors, even when allowed to peek into the solution
abstract
Large language models (LLMs) demonstrate remarkable performance on math word problems, yet they have been shown to struggle with meta-reasoning tasks such as identifying errors in student solutions.In this work, we investigate the challenge of locating the first error step in stepwise solutions using two error reasoning datasets: VtG and PRM800K.Our experiments show that state-of-the-art LLMs struggle to locate the first error step in student solutions even when given access to the reference solution.To that end, we propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student's solution, which helps improve performance.
KV Aditya Srivatsa, Kaushal Kumar Maurya, Ekaterina Kochmar
EMNLP3
2025 Generative AI for Early Grade Story Generation Using a Self-Reflective Approach
Taufiq Syed, Aadhith Shankarnarayanan, Yara Kaddoura, Salsabeel Y. Shapsough, Imran A. Zualkernan, Ekaterina Kochmar
ICALT6
2025 Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
abstract
Kaushal Kumar Maurya, Kv Aditya Srivatsa, Kseniia Petukhova, Ekaterina Kochmar. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kaushal Kumar Maurya, KV Aditya Srivatsa, Kseniia Petukhova, Ekaterina Kochmar
NAACL (Long Papers)4
2024 How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
abstract
Question generation (QG) is a natural language processing task with an abundance of potential benefits and use cases in the educational domain. In order for this potential to be realized, QG systems must be designed and validated with pedagogical needs in mind. However, little research has assessed or designed QG approaches with the input of real teachers or students. This paper applies a large language model-based QG approach where questions are generated with learning goals derived from Bloom's taxonomy. The automatically generated questions are used in multiple experiments designed to assess how teachers use them in practice. The results demonstrate that teachers prefer to write quizzes with automatically generated questions, and that such quizzes have no loss in quality compared to handwritten versions. Further, several metrics indicate that automatically generated questions can even improve the quality of the quizzes created, showing the promise for large scale use of QG in the classroom setting.
Sabina Elkins, Ekaterina Kochmar, Jackie Chi Kit Cheung, Iulian Serban
AAAI2
2024 REFeREE: A REference-FREE Model-Based Metric for Text Simplification
abstract
Text simplification lacks a universal standard of quality, and annotated reference simplifications are scarce and costly. We propose to alleviate such limitations by introducing REFeREE, a reference-free model-based metric with a 3-stage curriculum. REFeREE leverages an arbitrarily scalable pretraining stage and can be applied to any quality standard as long as a small number of human annotations are available. Our experiments show that our metric outperforms existing reference-based metrics in predicting overall ratings and reaches competitive and consistent performance in predicting specific ratings while requiring no reference simplifications at inference time.
Ekaterina Kochmar
LREC/COLING2
2023 BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine Languages
abstract
Current research on automatic readability assessment (ARA) has focused on improving the performance of models in high-resource languages such as English.In this work, we introduce and release BASAHACORPUS as part of an initiative aimed at expanding available corpora and baseline models for readability assessment in lower resource languages in the Philippines.We compiled a corpus of short fictional narratives written in Hiligaynon, Minasbate, Karaya, and Rinconada-languages belonging to the Central Philippine family tree subgroup-to train ARA models using surface-level, syllablepattern, and n-gram overlap features.We also propose a new hierarchical cross-lingual modeling approach that takes advantage of a language's placement in the family tree to increase the amount of available training data.Our study yields encouraging results that support previous work showcasing the efficacy of cross-lingual models in low-resource settings, as well as similarities in highly informative linguistic features for mutually intelligible languages.1
Joseph Marvin Imperial, Ekaterina Kochmar
EMNLP2
2022 Raising Student Completion Rates with Adaptive Curriculum and Contextual Bandits
Robert Belfer, Ekaterina Kochmar, Iulian Serban
AIED (1)2
2021 Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor Systems
abstract
We explore creating automated, personalized feedback in an intelligent tutoring system (ITS). Our goal is to pinpoint correct and incorrect concepts in student answers in order to achieve better student learning gains. Although automatic methods for providing personalized feedback exist, they do not explicitly inform students about which concepts in their answers are correct or incorrect. Our approach involves decomposing students answers using neural discourse segmentation and classification techniques. This decomposition yields a relational graph over all discourse units covered by the reference solutions and student answers. We use this inferred relational graph structure and a neural classifier to match student answers with reference solutions and generate personalized feedback. Although the process is completely automated and data-driven, the personalized feedback generated is highly contextual, domain-aware and effectively targets each student's misconceptions and knowledge gaps. We test our method in a dialogue-based ITS and demonstrate that our approach results in high-quality feedback and significantly improved student learning gains.
Matt Grenander, Robert Belfer, Ekaterina Kochmar, Iulian Serban, François St-Hilaire, Jackie Chi Kit Cheung
AAAI3
2021 A Comparative Study of Learning Outcomes for Online Learning Platforms
François St-Hilaire, Nathan Burns, Robert Belfer, Muhammad Shayan, Ariella Smofsky, Dung Do Vu, Antoine Frau, Joseph Potochny, Farid Faraji, Vincent Pavero, Neroli Ko, Ansona Onyi Ching, Sabina Elkins, Anush Stepanyan, Adela Matajova, Laurent Charlin, Yoshua Bengio, Iulian Serban, Ekaterina Kochmar
AIED (2)19
2021 A New Readability Assessment Tool
Rebecca Watson, Ekaterina Kochmar
EDM2
2021 Word Complexity is in the Eye of the Beholder
abstract
Sian Gooding, Ekaterina Kochmar, Seid Muhie Yimam, Chris Biemann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Sian Gooding, Ekaterina Kochmar, Seid Muhie Yimam, Chris Biemann
NAACL-HLT2
2020 Automated Personalized Feedback Improves Learning Gains in An Intelligent Tutoring System
Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Iulian Serban, Joelle Pineau
AIED (2)1
2020 A Large-Scale, Open-Domain, Mixed-Interface Dialogue-Based ITS for STEM
Iulian Serban, Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Joelle Pineau, Aaron C. Courville, Laurent Charlin, Yoshua Bengio
AIED (2)3
2020 Detecting Multiword Expression Type Helps Lexical Complexity Assessment
abstract
Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature. Multiple NLP applications have been shown to benefit from MWE identification, however the research on lexical complexity of MWEs is still an under-explored area. In this work, we re-annotate the Complex Word Identification Shared Task 2018 dataset of Yimam et al. (2017), which provides complexity scores for a range of lexemes, with the types of MWEs. We release the MWE-annotated dataset with this paper, and we believe this dataset represents a valuable resource for the text simplification community. In addition, we investigate which types of expressions are most problematic for native and non-native readers. Finally, we show that a lexical complexity assessment system benefits from the information about MWE types.
Ekaterina Kochmar, Sian Gooding, Matthew Shardlow
LREC1
2020 SeCoDa: Sense Complexity Dataset
abstract
The Sense Complexity Dataset (SeCoDa) provides a corpus that is annotated jointly for complexity and word senses. It thus provides a valuable resource for both word sense disambiguation and the task of complex word identification. The intention is that this dataset will be used to identify complexity at the level of word senses rather than word tokens. For word sense annotation SeCoDa uses a hierarchical scheme that is based on information available in the Cambridge Advanced Learner’s Dictionary. This way we can offer more coarse-grained senses than directly available in WordNet.
David Strohmaier, Sian Gooding, Shiva Taslimipoor, Ekaterina Kochmar
LREC4
2019 Complex Word Identification as a Sequence Labelling Task
abstract
Complex Word Identification (CWI) is concerned with detection of words in need of simplification and is a crucial first step in a simplification pipeline.It has been shown that reliable CWI systems considerably improve text simplification.However, most CWI systems to date address the task on a word-by-word basis, not taking the context into account.In this paper, we present a novel approach to CWI based on sequence modelling.Our system is capable of performing CWI in context, does not require extensive feature engineering and outperforms state-of-the-art systems on this task.
Sian Gooding, Ekaterina Kochmar
ACL (1)2
2019 Recursive Context-Aware Lexical Simplification
abstract
Sian Gooding, Ekaterina Kochmar. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sian Gooding, Ekaterina Kochmar
EMNLP/IJCNLP (1)2
2017 Classification of Twitter Accounts into Automated Agents and Human Users
abstract
Online social networks (OSNs) have seen a remarkable rise in the presence of surreptitious automated accounts. Massive human user-base and business-supportive operating model of social networks (such as Twitter) facilitates the creation of automated agents. In this paper we outline a systematic methodology and train a classifier to categorise Twitter accounts into 'automated' and 'human' users. To improve classification accuracy we employ a set of novel steps. First, we divide the dataset into four popularity bands to compensate for differences in types of accounts. Second, we create a large ground truth dataset using human annotations and extract relevant features from raw tweets. To judge accuracy of the procedure we calculate agreement among human annotators as well as with a bot detection research tool. We then apply a Random Forests classifier that achieves an accuracy close to human agreement. Finally, as a concluding step we perform tests to measure the efficacy of our results.
Zafar Gilani, Ekaterina Kochmar, Jon Crowcroft
ASONAM2
2016 Cross-Lingual Lexico-Semantic Transfer in Language Learning
abstract
Lexico-semantic knowledge of our native language provides an initial foundation for second language learning.In this paper, we investigate whether and to what extent the lexico-semantic models of the native language (L1) are transferred to the second language (L2).Specifically, we focus on the problem of lexical choice and investigate it in the context of three typologically diverse languages: Russian, Spanish and English.We show that a statistical semantic model learned from L1 data improves automatic error detection in L2 for the speakers of the respective L1.Finally, we investigate whether the semantic model learned from a particular L1 is portable to other, typologically related languages.
Ekaterina Kochmar, Ekaterina Shutova
ACL (1)1
2016 'Calling on the classical phone': a distributional model of adjective-noun errors in learners' English
abstract
In this paper we discuss three key points related to error detection (ED) in learners’ English. We focus on content word ED as one of the most challenging tasks in this area, illustrating our claims on adjective–noun (AN) combinations. In particular, we (1) investigate the role of context in accurately capturing semantic anomalies and implement a system based on distributional topic coherence, which achieves state-of-the-art accuracy on a standard test set; (2) thoroughly investigate our system’s performance across individual adjective classes, concluding that a class-dependent approach is beneficial to the task; (3) discuss the data size bottleneck in this area, and highlight the challenges of automatic error generation for content words.
Aurélie Herbelot, Ekaterina Kochmar
COLING2
2014 Detecting Learner Errors in the Choice of Content Words Using Compositional Distributional Semantics
Ekaterina Kochmar, Ted Briscoe
COLING1