VLDB 2026 Research / reviewers in the wild / expert
Steven Moore
dblp:174/0197
· DBLP profile ↗
24ranked-venue papers
16as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 15 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 13 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 7 first-author · 6 since 2021Systems, architecture and hardware · 7 · 7 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tools, Teammates, or Threats? How Pedagogical Reasoning Shapes Novice Instructional Designers' Judgments About AI in Education
Steven Moore, Christine Kwon |
AIED | 1 |
| 2025 | Generative AI in Instructional Design Education: Effects on Novice Microlesson Quality
Steven Moore, Lydia Eckstein, Christine Kwon, John C. Stamper |
AIED (4) | 1 |
| 2025 | Platform-based Adaptive Experimental Research in Education: Lessons Learned from The Digital Learning ChallengeabstractAdaptive Experimentation is one of the most promising approaches to support complex decision-making in learning experience design and delivery. This paper reports on our experience with a real-world, multi-experimental evaluation of an adaptive experimentation platform within the XPRIZE Digital Learning Challenge framework, and summarizes data-driven lessons learned and best practices for Adaptive Experimentation in education. We outline key scenarios of the applicability of platform-supported experiments and reflect on lessons learned from this two-year project, focusing on implications relevant to platform developers, researchers, practitioners, and policy stakeholders to integrate Adaptive Experiments in real-world courses. Ilya Musabirov, Mohi Reza, Haochen Song, Steven Moore, Pan Chen 0005, John C. Stamper, Norman L. Bier, Anna N. Rafferty, Thomas W. Price, Nina Deliu, Audrey Durand, Michael Liut, Joseph Jay Williams |
LAK | 4 |
| 2025 | Learnersourcing: Student-generated Content @ Scale: 3rd Annual WorkshopabstractPeer Reviewed Steven Moore, Xinyi Lu 0004, Hyoungwook Jin, Hassan Khosravi, Paul Denny 0001, Christopher Brooks 0001, Xu Wang 0016, Juho Kim 0001, John C. Stamper |
L@S | 1 |
| 2024 | An Automatic Question Usability Evaluation Toolkit
Steven Moore, Eamon Costello, Huy Anh Nguyen, John C. Stamper |
AIED (2) | 1 |
| 2024 | Leveraging Large Language Models for Next-Generation Educational Technologies
Neil T. Heffernan, Rose E. Wang, Christopher J. MacLellan, Arto Hellas, Chenglu Li, Candace A. Walkington, Joshua Littenberg-Tobias, David Joyner, Steven Moore, Adish Singla, Zachary A. Pardos, Maciej Pankiewicz, Juho Kim 0001, Shashank Sonkar, Clayton Cohn, Anthony Botelho, Andrew S. Lan, Mingyu Feng, Tanja Käser, Eamon Worden |
EDM | 9 |
| 2024 | Learnersourcing: Student-generated Content @ Scale: 2nd Annual Workshopabstractaendees to leave the workshop with a practical understanding of how to engage with learnersourcing.Participants will get hands-on experience with current tools, Steven Moore, Xinyi Lu 0004, Hyoungwook Jin, Hassan Khosravi, Paul Denny 0001, Christopher Brooks 0001, Xu Wang 0016, Juho Kim 0001, John C. Stamper |
L@S | 1 |
| 2024 | Automated Generation and Tagging of Knowledge Components from Multiple-Choice QuestionsabstractKnowledge Components (KCs) linked to assessments enhance the measurement of student learning, enrich analytics, and facilitate adaptivity. However, generating and linking KCs to assessment items requires significant effort and domain-specific knowledge. To streamline this process for higher-education courses, we employed GPT-4 to generate KCs for multiple-choice questions (MCQs) in Chemistry and E-Learning. We analyzed discrepancies between the KCs generated by the Large Language Model (LLM) and those made by humans through evaluation from three domain experts in each subject area. This evaluation aimed to determine whether, in instances of non-matching KCs, evaluators showed a preference for the LLM-generated KCs over their human-created counterparts. We also developed an ontology induction algorithm to cluster questions that assess similar KCs based on their content. Our most effective LLM strategy accurately matched KCs for 56% of Chemistry and 35% of E-Learning MCQs, with even higher success when considering the top five KC suggestions. Human evaluators favored LLM-generated KCs, choosing them over human-assigned ones approximately two-thirds of the time, a preference that was statistically significant across both domains. Our clustering algorithm successfully grouped questions by their underlying KCs without needing explicit labels or contextual information. This research advances the automation of KC generation and classification for assessment items, alleviating the need for student data or predefined KC labels. Steven Moore, Robin Schmucker, Tom M. Mitchell, John C. Stamper |
L@S | 1 |
| 2023 | fAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision TasksabstractTo design with AI models, user experience (UX) designers must assess the fit between the model and user needs. Based on user research, they need to contextualize the model’s behavior and potential failures within their product-specific data instances and user scenarios. However, our formative interviews with ten UX professionals revealed that such a proactive discovery of model limitations is challenging and time-intensive. Furthermore, designers often lack technical knowledge of AI and accessible exploration tools, which challenges their understanding of model capabilities and limitations. In this work, we introduced a failure-driven design approach to AI, a workflow that encourages designers to explore model behavior and failure patterns early in the design process. The implementation of fAIlureNotes, a designer-centered failure exploration and analysis tool, supports designers in evaluating models and identifying failures across diverse user groups and scenarios. Our evaluation with UX practitioners shows that fAIlureNotes outperforms today’s interactive model cards in assessing context-specific model performance. Steven Moore, Qingzi Vera Liao, Hariharan Subramonyam |
CHI | 1 |
| 2023 | Assessing the Quality of Multiple-Choice Questions Using GPT-4 and Rule-Based Methods
Steven Moore, Huy Anh Nguyen, Tianying Chen 0001, John C. Stamper |
EC-TEL | 1 |
| 2023 | Tracking Knowledge for Learning Japanese as a 2nd Language
Tomoko Okimoto, Matthew W. Johnson 0001, Huy Anh Nguyen, Steven Moore, Michael Eagle, John C. Stamper |
ICCE | 4 |
| 2023 | Crowdsourcing the Evaluation of Multiple-Choice Questions Using Item-Writing Flaws and Bloom's TaxonomyabstractMultiple-choice questions, which are widely used in educational assessments, have the potential to negatively impact student learning and skew analytics when they contain item-writing flaws. Existing methods for evaluating multiple-choice questions in educational contexts tend to focus primarily on machine readability metrics, such as grammar, syntax, and formatting, without considering the intended use of the questions within course materials and their pedagogical implications. In this study, we present the results of crowdsourcing the evaluation of multiple-choice questions based on 15 common item-writing flaws. Through analysis of 80 crowdsourced evaluations on questions from the domains of calculus and chemistry, we found that crowdworkers were able to accurately evaluate the questions, matching 75% of the expert evaluations across multiple questions. They were able to correctly distinguish between two levels of Bloom's Taxonomy for the calculus questions, but were less accurate for chemistry questions. We discuss how to scale this question evaluation process and the implications it has across other domains. This work demonstrates how crowdworkers can be leveraged in the quality evaluation of educational questions, regardless of prior experience or domain knowledge. Steven Moore, Ellen Fang, Huy Anh Nguyen, John C. Stamper |
L@S | 1 |
| 2022 | Assessing the Quality of Student-Generated Short Answer Questions Using GPT-3
Steven Moore, Huy Anh Nguyen, Norman L. Bier, Tanvi Domadia, John C. Stamper |
EC-TEL | 1 |
| 2022 | Towards Generalized Methods for Automatic Question Generation in Educational Domains
Huy Anh Nguyen, Shravya Bhat, Steven Moore, Norman L. Bier, John C. Stamper |
EC-TEL | 3 |
| 2022 | Towards Automated Generation and Evaluation of Questions in Educational Domains
Shravya Bhat, Huy Anh Nguyen, Steven Moore, John C. Stamper, Majd F. Sakr, Eric Nyberg |
EDM | 3 |
| 2022 | Learnersourcing: Student-generated Content @ ScaleabstractThe first annual workshop on Learnersourcing: Student-generated Content @ Scale is taking place at Learning @ Scale 2022. This hybrid workshop will expose attendees to the ample opportunities in the learnersourcing space, including instructors, researchers, learning engineers, and many other roles. We believe participants from a wide range of backgrounds and prior knowledge on learnersourcing can both benefit and contribute to this workshop, as learnersourcing draws on work from education, crowdsourcing, learning analytics, data mining, ML/NLP, and many more fields. Additionally, as the learnersourcing process involves many stakeholders (students, instructors, researchers, instructional designers, etc.), multiple viewpoints can help to inform what future student-generated content might be useful, new and better ways to assess the quality of the content and spark potential collaboration efforts between attendees. We ultimately want to show how everyone can make use of learnersourcing and have participants gain hands-on experience using existing tools, create their own learnersourcing activities using them or their own platforms, and take part in discussing the next challenges and opportunities in the learnersourcing space. Our hope is to attract attendees interested in scaling the generation of instructional and assessment content and those interested in the use of online learning platforms. Steven Moore, John C. Stamper, Christopher Brooks 0001, Paul Denny 0001, Hassan Khosravi |
L@S | 1 |
| 2021 | Exploring Metrics for the Analysis of Code Submissions in an Introductory Data Science CourseabstractWhile data science education has gained increased recognition in both academic institutions and industry, there has been a lack of research on automated coding assessment for novice students. Our work presents a first step in this direction, by leveraging the coding metrics from traditional software engineering (Halstead Volume and Cyclomatic Complexity) in combination with those that reflect a data science project’s learning objectives (number of library calls and number of common library calls with the solution code). Through these metrics, we examined the code submissions of 97 students across two semesters of an introductory data science course. Our results indicated that the metrics can identify cases where students had overly complicated codes and would benefit from scaffolding feedback. The number of library calls, in particular, was also a significant predictor of changes in submission score and submission runtime, which highlights the distinctive nature of data science programming. We conclude with suggestions for extending our analyses towards more actionable intervention strategies, for example by tracking the fine-grained submission grading outputs throughout a student’s submission history, to better model and support them in their data science learning process. Huy Anh Nguyen, Michelle Lim, Steven Moore, Eric Nyberg, Majd F. Sakr, John C. Stamper |
LAK | 3 |
| 2021 | Examining the Effects of Student Participation and Performance on the Quality of Learnersourcing Multiple-Choice QuestionsabstractWhile generating multiple-choice questions has been shown to promote deep learning, students often fail to realize this benefit and do not willingly participate in this activity. Additionally, the quality of the student-generated questions may be influenced by both their level of engagement and familiarity with the learning materials. Towards better understanding how students can generate high quality questions, we designed and deployed a multiple-choice question generation activity in seven college-level online chemistry courses. From these courses, we collected data on student interactions and their contribution to the question-generation task. A total of 201 students enrolled in the courses and 57 of them elected to generate a multiple-choice question. Our results indicated that students were able to contribute quality questions, with 67% of them being evaluated by experts as acceptable for use. We further identified several student behaviors in the online courses that are correlated to their participation in the task and the quality of their contribution. Our findings can help teachers and students better understand the benefits of student-generated questions and effectively implement future learnersourcing activities. Steven Moore, Huy Anh Nguyen, John C. Stamper |
L@S | 1 |
| 2020 | Evaluating Crowdsourcing and Topic Modeling in Generating Knowledge Components from Explanations
Steven Moore, Huy Anh Nguyen, John C. Stamper |
AIED (1) | 1 |
| 2020 | Utilizing Crowdsourcing and Topic Modeling to Generate Knowledge Components for Math and Writing Problems
Steven Moore, Huy Nugyen, John C. Stamper |
ICCE | 1 |
| 2020 | Towards Crowdsourcing the Identification of Knowledge ComponentsabstractAssigning a set of hypothesized knowledge components (KCs) to assessment items within an ed-tech system enables us to better estimate student learning. However, creating and assigning these KCs is a time-consuming process that often requires domain expertise. In this study, we present the results of crowdsourcing KCs for problems in the domain of mathematics and English writing, as a first step in leveraging the crowd to expedite this task. Crowdworkers were presented with a problem and asked to provide the underlying skills required to solve it. Additionally, we investigated the effect of priming crowdworkers with related content before having them generate these KCs. We then analyzed their contributions through qualitative coding and found that across both the math and writing domains roughly 33% of the crowdsourced KCs directly matched those generated by domain experts for the same problems. Steven Moore, Huy Anh Nguyen, John C. Stamper |
L@S | 1 |
| 2019 | Exploring Teachable Humans and Teachable Agents: Human Strategies Versus Agent Policies and the Basis of Expertise
John C. Stamper, Steven Moore |
AIED (2) | 2 |
| 2019 | Decision Support for an Adversarial Game Environment Using Automatic Hint Generation
Steven Moore, John C. Stamper |
ITS | 1 |
| 2017 | Moonstone: Support for understanding and writing exception handling codeabstractMoonstone is a new plugin for Eclipse that supports developers in understanding exception flow and in writing exception handlers in Java. Understanding exception control flow is paramount for writing robust exception handlers, a task many developers struggle with. To help with this understanding, we present two new kinds of information: ghost comments, which are transient overlays that reveal potential sources of exceptions directly in code, and annotated highlights of skipped code and associated handlers. To help developers write better handlers, Moonstone additionally provides project-specific recommendations, detects common bad practices, such as empty or inadequate handlers, and provides automatic resolutions, introducing programmers to advanced Java exception handling features, such as try-with-resources. We present findings from two formative studies that informed the design of Moonstone. We then show with a user study that Moonstone improves users' understanding in certain areas and enables developers to amend exception handling code more quickly and correctly. Florian Kistner, Mary Beth Kery, Michael Puskas, Steven Moore, Brad A. Myers |
VL/HCC | 4 |