Owen Henkel

dblp:348/0438 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0001-8850-067XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work
Owen Henkel, Bill Roberts, Doug Jaffe, Laurence Holt
AIED1
2025 How Much Mastery is Enough Mastery? The Relationship between Mastery in a Lesson and the Performance on the Subsequent Lesson
Jiayi Zhang 0004, Kirk Vanacore, Ryan Baker 0001, Nabil Ch, Caitlin Mills 0001, Owen Henkel
EDM6
2024 Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference
Owen Henkel, Zachary Levonian, Chenglu Li, Millie-Ellen Postle
EDM1
2024 Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability To Mark Short Answer Questions in K-12 Education
abstract
This paper presents reports on a series of experiments with a novel dataset evaluating how well Large Language Models (LLMs) can mark (i.e. grade) open text responses to short answer questions, Specifically, we explore how well different combinations of GPT version and prompt engineering strategies performed at marking real student answers to short answer across different domain areas (Science and History) and grade-levels (spanning ages 5-16) using a new, never-used-before dataset from Carousel, a quizzing platform. We found that GPT-4, with basic few-shot prompting performed well (Kappa, 0.70) and, importantly, very close to human-level performance (0.75). This research builds on prior findings that GPT-4 could reliably score short answer reading comprehension questions at a performance-level very close to that of expert human raters. The proximity to human-level performance, across a variety of subjects and grade levels suggests that LLMs could be a valuable tool for supporting low-stakes formative assessment tasks in K-12 education and has important implications for real-world education delivery.
Owen Henkel, Libby Hills, Adam Boxer, Bill Roberts, Zachary Levonian
L@S1
2023 Leveraging Human Feedback to Scale Educational Datasets: Combining Crowdworkers and Comparative Judgement
abstract
Machine Learning models have many potentially beneficial applications in education settings, but a key barrier to their development is securing enough high-quality, labelled data to train these models. This process has traditionally relied on highly skilled raters using complex, multi-class rubrics, which made labelling expensive and difficult to scale. A more scalable approach would be to use non-expert crowdworkers to evaluate student work, but maintaining high levels of accuracy and inter-rater reliability when using non-expert workers can be challenging. This paper reports on two experiments in which non-expert crowdworkers hired to evaluate (i.e., score) student work and were randomly assigned to one of two conditions: the control, where they were asked to assign a rubrics based score (i.e., a categorical judgement), or the treatment, where they were shown the same student answers, but were asked to decide which of two candidate answers was better (i.e., a comparative/preference-based judgement). We found that using comparative judgement substantially improved inter-rater reliability on both tasks. These results are in-line with well-established literature on the benefits of comparative judgement in the field of educational assessment, as well as with recent trends in artificial intelligence research, where comparative judgement is becoming the preferred method for providing human feedback on model outputs. These results are novel and important in demonstrating the effects of using the combination of comparative judgement and crowdworkers to evaluate educational data
Owen Henkel, Libby Hills
L@S1