Detmar Meurers

dblp:m/WaltDetmarMeurers · also Walt Detmar Meurers · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-9740-7442ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Who Did What to Succeed? Individual Differences in Which Learning Behaviors Are Linked to Achievement
abstract
It is commonly assumed that digital learning environments such as intelligent tutoring systems facilitate learning and positively impact achievement. This study explores how different groups of students exhibit distinct relationships between learning behaviors and academic achievement in an intelligent tutoring system for English as a foreign language. We examined whether these differences are linked to students’ prior knowledge, personality traits, and motivation. We collected behavioral trace data from 507 German seventh-grade students during the 2021/22 school year and applied machine learning models to predict English performance based on learning behaviors (best-performing model’s 2 = .41). To understand the impact of specific behaviors, we applied the explainable AI method SHAP and identified three student clusters with distinct learning behavior patterns. Subsequent analyses revealed that these clusters also varied in prior knowledge and motivation: one with high prior knowledge and average motivation, another with low prior knowledge and average motivation, and a third with both low prior knowledge and low motivation. Our findings suggest that learning behaviors are linked differently to academic success across students and are closely tied to their prior knowledge and motivation. This hints towards the importance of personalizing learning systems to support individual learning needs better.
Hannah Deininger, Cora Parrisius, Rosa Lavelle-Hill, Detmar Meurers, Ulrich Trautwein, Benjamin Nagengast, Gjergji Kasneci
LAK4
2025 Grammar Control in Dialogue Response Generation for Language Learning Chatbots
abstract
Dominik Glandorf, Peng Cui, Detmar Meurers, Mrinmaya Sachan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Dominik Glandorf, Peng Cui 0006, Detmar Meurers, Mrinmaya Sachan
NAACL (Long Papers)3
2024 Towards Task-Oriented ICALL: A Criterion-Referenced Learner Dashboard Organising Digital Practice
Leona Colling, Ines Pieronczyk, Cora Parrisius, Heiko Holz, Stephen Bodnar, Florian Nuxoll, Detmar Meurers
CSEDU (1)7
2023 Can You Solve This on the First Try? - Understanding Exercise Field Performance in an Intelligent Tutoring System
Hannah Deininger, Rosa Lavelle-Hill, Cora Parrisius, Ines Pieronczyk, Leona Colling, Detmar Meurers, Ulrich Trautwein, Benjamin Nagengast, Gjergji Kasneci
AIED6
2021 Exploring Input Representation Granularity for Generating Questions Satisfying Question-Answer Congruence
abstract
In question generation, the question produced has to be well-formed and meaningfully related to the answer serving as input.Neural generation methods have predominantly leveraged the distributional semantics of words as representations of meaning and generated questions one word at a time.In this paper, we explore the viability of form-based and more fine-grained encodings such as character or subword representations for question generation.We start from the typical seq2seq architecture using word embeddings presented by De Kuthy et al. (2020), who generate questions from text so that the answer given in the input text matches not just in meaning but also in form, satisfying question-answer congruence.We show that models trained on character and subword representations substantially outperform the published results based on word embeddings, and they do so with fewer parameters.Our approach eliminates two important problems of the word-based approach: the encoding of rare or out-of-vocabulary words and the incorrect replacement of words with semantically-related ones.The characterbased model substantially improves on the published results, both in terms of BLEU scores and regarding the quality of the generated question.Going beyond the specific task, this result adds to the evidence weighing different form-and meaning-based representations for natural language processing tasks.
Madeeswaran Kannan, Haemanth Santhi Ponnusamy, Kordula De Kuthy, Lukas Stein, Detmar Meurers
INLG5
2021 Development and Evaluation of a Tablet-Based Reading Fluency Test for Primary School Children
abstract
Assessing children’s literacy skills is a key requirement for successful learning. However, standardized assessments are almost exclusively available as paper-and-pencil tests, discarding digital testing’s advantages. In this article, we develop and evaluate two different tablet versions for a paper-based reading fluency test for German primary school children: one true-to-original design (TDId) and a tablet-optimized design (TDOpt). We investigate the reliability, equivalence, and user experience of the tablet-based versions by comparing it to each other and to the paper-based version in a user test with 21 fourth graders. The results suggest very high reliability of both tablet-based versions (r’s >.94). Children scored significantly higher on TDId than on the other tests while reading scores assessed with TDOpt matched the conventional test. Children preferred both tablet versions over the paper-based test, whereby TDOpt was consistently rated best. This study suggests that a tablet-optimized reading fluency test not only retains test reliability but is also more appealing to primary school children.
Kristina Dawidowsky, Heiko Holz, Jakob Schwerter, Ines Pieronczyk, Detmar Meurers
MobileHCI5
2021 Interaction Styles in Context: Comparing Drag-and-Drop, Point-and-Touch, and Touch in a Mobile Spelling Game
abstract
The choice of appropriate interaction style for children’s application has been a divisive subject of debate. For example, drag-and-drop is often said to perform worse than point-and-click in educational applications – and vice versa. In this article, we argue for the need to choose the interaction style in context, considering a range of factors. We compare drag-and-drop, point-and-touch, and simple touch for selecting letters to form words in a spelling line as part of an educational spelling game. We evaluate the perceived workload, user experience, preference, and writing times of twenty-five children (8–11 years), eight of whom were dyslexic. We found that touch received better ratings and was ranked highest most often on all subscales compared to drag-and-drop and point-and-touch. Children needed less time using touch and 68% chose it as their favorite interaction style. We also found small advantages for drag-and-drop over point-and-touch, which runs counter to some recent recommendations. This becomes particularly clear when using ranking responses, which support a particularly fine-grained picture.
Heiko Holz, Detmar Meurers
Int. J. Hum. Comput. Interact.2
2021 Toward neuroadaptive support technologies for improving digital reading: a passive BCI-based assessment of mental workload imposed by text difficulty and presentation speed during reading
abstract
Abstract We investigated whether a passive brain–computer interface that was trained to distinguish low and high mental workload in the electroencephalogram (EEG) can be used to identify (1) texts of different readability difficulties and (2) texts read at different presentation speeds. For twelve subjects we calibrated a subject-dependent, but task-independent predictive model classifying mental workload. We then recorded EEG data from each subject, while twelve texts in blocks of three were presented to them word by word. Half of the texts were easy, and the other half were difficult texts according to classic reading formulas. From each text category three texts were read at a self-adjusted comfortable presentation speed and the other three at an increased speed. For each subject we applied the predictive model to EEG data of each word of the twelve texts. We found that the resulting predictive values for mental workload were higher for difficult texts than for easy texts. Predictive values from texts presented at an increased speed were also higher than for those presented at a normal self-adjusted speed. The results suggest that the task-independent predictive model can be used on single-subject level to build a highly predictive user model of the reader over time. Such a model could be employed in a system which continuously monitors brain activity related to mental workload and adapts to specific reader’s abilities and characteristics by adjusting the difficulty of text materials and the way it is presented to the reader in real time. A neuroadaptive system like this could foster efficient reading and text-based learning by keeping readers’ mental workload levels at an individually optimal level.
Lena M. Andreessen, Peter Gerjets, Detmar Meurers, Thorsten O. Zander
User Model. User Adapt. Interact.3
2020 Towards automatically generating Questions under Discussion to link information and discourse structure
abstract
Questions under Discussion (QUD; Roberts, 2012) are emerging as a conceptually fruitful approach to spelling out the connection between the information structure of a sentence and the nature of the discourse in which the sentence can function.To make this approach useful for analyzing authentic data, Riester, Brunetti & De Kuthy (2018) presented a discourse annotation framework based on explicit pragmatic principles for determining a QUD for every assertion in a text.De Kuthy et al. (2018) demonstrate that this supports more reliable discourse structure annotation, and Ziai and Meurers (2018) show that based on explicit questions, automatic focus annotation becomes feasible.But both approaches are based on manually specified questions.In this paper, we present an automatic question generation approach to partially automate QUD annotation by generating all potentially relevant questions for a given sentence.While transformation rules can concisely capture the typical question formation process, a rule-based approach is not sufficiently robust for authentic data.We therefore employ the transformation rules to generate a large set of sentence-question-answer triples and train a neural question generation model on them to obtain both systematic question type coverage and robustness.
Kordula De Kuthy, Madeeswaran Kannan, Haemanth Santhi Ponnusamy, Detmar Meurers
COLING4
2018 Modeling the Readability of German Targeting Adults and Children: An empirically broad analysis and its cross-corpus validation
abstract
We analyze two novel data sets of German educational media texts targeting adults and children. The analysis is based on 400 automatically extracted measures of linguistic complexity from a wide range of linguistic domains. We show that both data sets exhibit broad linguistic adaptation to the target audience, which generalizes across both data sets. Our most successful binary classification model for German readability robustly shows high accuracy between 89.4%–98.9% for both data sets. To our knowledge, this comprehensive German readability model is the first for which robust cross-corpus performance has been shown. The research also contributes resources for German readability assessment that are externally validated as successful for different target audiences: we compiled a new corpus of German news broadcast subtitles, the Tagesschau/Logo corpus, and crawled a GEO/GEOlino corpus substantially enlarging the data compiled by Hancke et al. 2012.
Zarah Weiß, Detmar Meurers
COLING2
2018 Automatic Focus Annotation: Bringing Formal Pragmatics Alive in Analyzing the Information Structure of Authentic Data
abstract
Ramon Ziai, Detmar Meurers. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Ramon Ziai, Detmar Meurers
NAACL-HLT2
2016 Advancing Linguistic Features and Insights by Label-informed Feature Grouping: An Exploration in the Context of Native Language Identification
abstract
We propose a hierarchical clustering approach designed to group linguistic features for supervised machine learning that is inspired by variationist linguistics. The method makes it possible to abstract away from the individual feature occurrences by grouping features together that behave alike with respect to the target class, thus providing a new, more general perspective on the data. On the one hand, it reduces data sparsity, leading to quantitative performance gains. On the other, it supports the formation and evaluation of hypotheses about individual choices of linguistic structures. We explore the method using features based on verb subcategorization information and evaluate the approach in the context of the Native Language Identification (NLI) task.
Serhiy Bykh, Detmar Meurers
COLING2
2016 Focus Annotation of Task-based Data: A Comparison of Expert and Crowd-Sourced Annotation in a Reading Comprehension Corpus
Kordula De Kuthy, Ramon Ziai, Detmar Meurers
LREC3
2014 Exploring Syntactic Features for Native Language Identification: A Variationist Perspective on Feature Encoding and Ensemble Optimization
Serhiy Bykh, Detmar Meurers
COLING2
2014 Assessing the relative reading level of sentence pairs for text simplification
abstract
While the automatic analysis of the readability of texts has a long history, the use of readability assessment for text simplification has received only little attention so far.In this paper, we explore readability models for identifying differences in the reading levels of simplified and unsimplified versions of sentences.Our experiments show that a relative ranking is preferable to an absolute binary one and that the accuracy of identifying relative simplification depends on the initial reading level of the unsimplified version.The approach is particularly successful in classifying the relative reading level of harder sentences.In terms of practical relevance, the approach promises to be useful for identifying particularly relevant targets for simplification and to evaluate simplifications given specific readability constraints.
Sowmya Vajjala, Detmar Meurers
EACL2
2014 The MERLIN corpus: Learner language and the CEFR
Adriane Boyd, Jirka Hana, Lionel Nicolas, Detmar Meurers, Katrin Wisniewski, Andrea Abel, Karin Schöne, Barbora Stindlová, Chiara Vettori
LREC4
2014 CLARA: A New Generation of Researchers in Common Language Resources and Their Applications
Koenraad De Smedt, Erhard W. Hinrichs, Detmar Meurers, Inguna Skadina, Bolette S. Pedersen, Costanza Navarretta, Núria Bel, Krister Lindén, Markéta Lopatková, Jan Hajic 0001, Gisle Andersen, Przemyslaw Lenkiewicz
LREC3
2012 Native Language Identification using Recurring n-grams - Investigating Abstraction and Domain Dependence
Serhiy Bykh, Detmar Meurers
COLING2
2012 Readability Classification for German using Lexical, Syntactic, and Morphological Features
Julia Hancke, Sowmya Vajjala, Detmar Meurers
COLING3
2005 Detecting Errors in Discontinuous Structural Annotation
abstract
Consistency of corpus annotation is an essential property for the many uses of annotated corpora in computational and theoretical linguistics. While some research addresses the detection of inconsistencies in positional annotation (e.g., part-of-speech) and continuous structural annotation (e.g., syntactic constituency), no approach has yet been developed for automatically detecting annotation errors in discontinuous structural annotation. This is significant since the annotation of potentially discontinuous stretches of material is increasingly relevant, from tree-banks for free-word order languages to semantic and discourse annotation.In this paper we discuss how the variation n-gram error detection approach (Dickinson and Meurers, 2003a) can be extended to discontinuous structural annotation. We exemplify the approach by showing how it successfully detects errors in the syntactic annotation of the German TIGER corpus (Brants et al., 2002).
Markus Dickinson, Detmar Meurers
ACL2
2004 A Grammar Formalism and Parser for Linearization-based HPSG
Michael W. Daniels, Detmar Meurers
COLING2
2003 Detecting Errors in Part-of-Speech Annotation
Markus Dickinson, Detmar Meurers
EACL2
1997 Interleaving Universal Principles and Relational Constraints over Typed Feature Logic
abstract
We introduce a typed feature logic system providing both universal implicational principles as well as definite clauses over feature terms. We show that such an architecture supports a modular encoding of linguistic theories and allows for a compact representation using underspecification. The system is fully implemented and has been used as a workbench to develop and test large HPSG grammars. The techniques described in this paper are not restricted to a specific implementation, but could be added to many current feature-based grammar development systems.
Thilo Götz, Detmar Meurers
ACL2
1997 A Computational Treatment of Lexical Rules in HPSG as Covariation in Lexical Entries
Detmar Meurers, Guido Minnen
Comput. Linguistics1
1995 Compiling HPSG Type Constraints into Definite Clause Programs
abstract
We present a new approach to HPSG processing: compiling HPSG grammars expressed as type constraints into definite clause programs. This provides a clear and computationally useful correspondence between linguistic theories and their implementation. The compiler performs offline constraint inheritance and code optimization. As a result, we are able to efficiently process with HPSG grammars without having to hand-translate them into definite clause or phrase structure based systems.
Thilo Götz, Detmar Meurers
ACL2