Anna N. Rafferty

dblp:26/6934 · DBLP profile ↗
← Back
44ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-8319-5370ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 14 · 5 first-author · 9 since 2021Systems, architecture and hardware · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Likelihood-Based Diagnosis with Generative Models: Confidence-Aware Measurement from Student Writing
Thomas Christie, Matthew Zent, Markus Hauru, Anna N. Rafferty, Simon Woodhead 0002
AIED (3)4
2026 Maximizing the Margin Between Desirable and Undesirable Elements in a Covering Problem
Sophie Boileau, Andrew Hong, David Liben-Nowell, Alistair Pattison, Anna N. Rafferty, Charlie Roslansky
COCOON5
2025 An Agentic Framework for Real-Time Pedagogical Plot Generation
Thomas Christie, Anna N. Rafferty, Zack Lee, Ella Cutler, Husni Almoubayyed
AIED (5)2
2025 Knowledge of Examples Affects Conditional Reasoning About Math
David W. Braithwaite, Anna N. Rafferty
CogSci2
2025 Just Read the Question: Enabling Generalization to New Assessment Items with Text Awareness
Arisha Khan, Nathaniel Li, Tori Shen, Anna N. Rafferty
EDM4
2025 Platform-based Adaptive Experimental Research in Education: Lessons Learned from The Digital Learning Challenge
abstract
Adaptive Experimentation is one of the most promising approaches to support complex decision-making in learning experience design and delivery. This paper reports on our experience with a real-world, multi-experimental evaluation of an adaptive experimentation platform within the XPRIZE Digital Learning Challenge framework, and summarizes data-driven lessons learned and best practices for Adaptive Experimentation in education. We outline key scenarios of the applicability of platform-supported experiments and reflect on lessons learned from this two-year project, focusing on implications relevant to platform developers, researchers, practitioners, and policy stakeholders to integrate Adaptive Experiments in real-world courses.
Ilya Musabirov, Mohi Reza, Haochen Song, Steven Moore, Pan Chen 0005, John C. Stamper, Norman L. Bier, Anna N. Rafferty, Thomas W. Price, Nina Deliu, Audrey Durand, Michael Liut, Joseph Jay Williams
LAK10
2024 Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental Health
abstract
Digital mental health (DMH) interventions, such as text-message-based lessons and activities, offer immense potential for accessible mental health support. While these interventions can be effective, real-world experimental testing can further enhance their design and impact. Adaptive experimentation, utilizing algorithms like Thompson Sampling for (contextual) multi-armed bandit (MAB) problems, can lead to continuous improvement and personalization. However, it remains unclear when these algorithms can simultaneously increase user experience rewards and facilitate appropriate data collection for social-behavioral scientists to analyze with sufficient statistical confidence. Although a growing body of research addresses the practical and statistical aspects of MAB and other adaptive algorithms, further exploration is needed to assess their impact across diverse real-world contexts. This paper presents a software system developed over two years that allows text-messaging intervention components to be adapted using bandit and other algorithms while collecting data for side-by-side comparison with traditional uniform random non-adaptive experiments. We evaluate the system by deploying a text-message-based DMH intervention to 1100 users, recruited through a large mental health non-profit organization, and share the path forward for deploying this system at scale. This system not only enables applications in mental health but could also serve as a model testbed for adaptive experimentation algorithms in other domains.
Jiakai Shi, Ilya Musabirov, Rachel Kornfield, Jonah Meyerhoff, Ananya Bhattacharjee, Chris J. Karr, Theresa Nguyen, David C. Mohr, Anna N. Rafferty, Sofia S. Villar, Nina Deliu, Joseph Jay Williams
AAAI11
2024 Uncertainty-preserving deep knowledge tracing with state-space models
Thomas Christie, Carson Cook, Anna N. Rafferty
EDM3
2024 Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in Classrooms
abstract
Self-reflection on learning experiences constitutes a fundamental cognitive process, essential for consolidating knowledge and enhancing learning efficacy. However, traditional methods to facilitate reflection often face challenges in personalization, immediacy of feedback, engagement, and scalability. Integration of Large Language Models (LLMs) into the reflection process could mitigate these limitations. In this paper, we conducted two randomized field experiments in undergraduate computer science courses to investigate the potential of LLMs to help students engage in post-lesson reflection. In the first experiment (N=145), students completed a take-home assignment with the support of an LLM assistant; half of these students were then provided access to an LLM designed to facilitate self-reflection. The results indicated that the students assigned to LLM-guided reflection reported somewhat increased self-confidence compared to peers in a no-reflection control and a non-significant trend towards higher scores on a later assessment. Thematic analysis of students' interactions with the LLM showed that the LLM often affirmed the student's understanding, expanded on the student's reflection, and prompted additional reflection; these behaviors suggest ways LLM-interaction might facilitate reflection. In the second experiment (N=112), we evaluated the impact of LLM-guided self-reflection against other scalable reflection methods, such as questionnaire-based activities and review of key lecture slides, after assignment. Our findings suggest that the students in the questionnaire and LLM-based reflection groups performed equally well and better than those who were only exposed to lecture slides, according to their scores on a proctored exam two weeks later on the same subject matter. These results underscore the utility of LLM-guided reflection and questionnaire-based activities in improving learning outcomes. Our work highlights that focusing solely on the accuracy of LLMs can overlook their potential to enhance metacognitive skills through practices such as self-reflection. We discuss the implications of our research for the learning-at-scale community, highlighting the potential of LLMs to enhance learning experiences through personalized, engaging, and scalable reflection practices.
Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Huayin Luo, Joseph Jay Williams, Anna N. Rafferty, John C. Stamper, Michael Liut
L@S9
2024 Growth in Knowledge of Programming Patterns: A Comparison Study of CS1 vs. CS2 Students
abstract
How does students' knowledge of code structure improve as they progress through their degree, and where do students struggle? We conducted a comparative study between introductory (CS1) and intermediate CS students (CS2) to explore these questions. Using an online survey with several tasks, including identification of expert patterns, judgment of readable structure, code comprehension, code writing, and editing, we focused on two important code structures: (S1) returning boolean expressions directly and (S2) unique vs. repeated code within if and else. Student performance varied based on structure and task: in both S1 and S2, CS2 students demonstrated higher performance in identifying patterns, judgment of readable structure, and editing. However, evidence of improvement in code writing was only found for S1, and improvement in code comprehension was only found for S2. Therefore, students may need different supports across different code structures. With the exception of comprehension of S1, student performance was far below ceiling, suggesting a need for more support.
Sara Nurollahian, Anna N. Rafferty, Noelle Brown, Eliane Wiese
SIGCSE (1)2
2024 Playing with Matches: Adopting Gale-Shapley for Managing Student Enrollments Beyond CS2
abstract
Enrollment in computer science has increased dramatically in recent years, straining capacities and leading to various strategies for managing enrollment. But some strategies increase student competition and may have disproportionate negative impacts on students from underrepresented groups. We believe success in computing education necessitates a more equitable approach to course enrollment. In this experience report, we describe our new enrollment mechanism, "the Match." Building on the Gale--Shapley stable matching algorithm, the Match was designed to encourage a liberal arts approach to course selection and attempt to broaden participation in computing. Drawing on data from three years of use, we find high student participation, with the vast majority of students having their enrollment preferences met. With Match registration, our courses have tended to be a bit more inclusive of younger students. The Match appears not to have disparate negative impacts like those of competitive enrollment, but has increased workload in the Registrar's Office. Overall, we believe the Match has decreased student and faculty angst around registration, and we argue that systems like the Match can help manage enrollment pressures in ways that are consistent with educational values.
Anna N. Rafferty, David Liben-Nowell, David R. Musicant, Emy Farley, Allie Lyman, Ann May
SIGCSE (1)1
2023 LENS: Predictive Diagnostics for Flexible and Efficient Assessments
abstract
The utility of assessment systems lies in their capacity to transform observations of student behavior into meaningful inferences about learning, knowledge, and skills. Common practice is to use latent variable models and produce scores on scales. However the simplicity of these psychometric models may filter out potentially valuable information present in student behavior. In particular, scale scores are not optimized to support granular instructional decisions. Machine learning offers promising alternatives, but proposed deep learning architectures are not ideally suited for operational testing conditions involving sparse data and shifting category labels for test questions.
Thomas Christie, Hayden Johnson, Carson Cook, Garron Gianopulos, Anna N. Rafferty
L@S5
2022 FATED 2022: Fairness, Accountability, and Transparency in Educational Data
Collin F. Lynch, Mirko Marras, Mykola Pechenizkiy, Anna N. Rafferty, Steven Ritter 0001, Vinitra Swamy, Renzhe Yu
EDM4
2022 Adversarial bandits for drawing generalizable conclusions in non-adversarial experiments: an empirical study
Zhi-Han Yang, Anna N. Rafferty
EDM3
2022 How can Email Interventions Increase Students' Completion of Online Homework? A Case Study Using A/B Comparisons
abstract
Email communication between instructors and students is ubiquitous, and it could be valuable to explore ways of testing out how to make email messages more impactful. This paper explores the design space of using emails to get students to plan and reflect on starting weekly homework earlier. We deployed a series of email reminders using randomized A/B comparisons to test alternative factors in the design of these emails, providing examples of an experimental paradigm and metrics for a broader range of interventions. We also surveyed and interviewed instructors and students to compare their predictions about the effectiveness of the reminders with their actual impact. We present our results on which seemingly obvious predictions about effective emails are not borne out, despite there being evidence for further exploring these interventions, as they can sometimes motivate students to attempt their homework more often. We also present qualitative evidence about student opinions and behaviours after receiving the emails, to guide further interventions. These findings provide insight into how to use randomized A/B comparisons in everyday channels such as emails, to provide empirical evidence to test our beliefs about the effectiveness of alternative design choices.
Angela M. Zavaleta Bernuy, Ziwen Han, Hammad Shaikh, Qi Yin Zheng, Lisa-Angelique Lim, Anna N. Rafferty, Andrew Petersen 0001, Joseph Jay Williams
LAK6
2022 Student Motivations and Goals for CS1: Themes and Variations
abstract
Students come to CS1 with a wide variety of motivations and goals, which may differ across subpopulations and be indicative of their future engagement with CS. While there is a rich literature relating success in CS1 to specific constructs, such as belonging, goal-orientation, or self-efficacy, less work has examined what motivations and goals students volunteer as most important for their enrollment in CS1. Here, we use qualitative coding to identify themes from students' open-ended descriptions of why they're taking CS1 and what they hope to get out of it, collected across fifteen years. Using quantitative analysis of these coded descriptions, and word-frequency analysis, we identify and name three clusters of students that encompass the majority of students taking CS1: Explorers, Planners, and Utilitarians. We also identify motivations and goals that are more common for particular populations, such as students who have not yet declared a major or students without prior programming experience, as well as factors predicting students' later engagement with CS. This work demonstrates the potential of qualitative coding and computational analyses to enable us to better understand a population of students based on their own words.
David Liben-Nowell, Anna N. Rafferty
SIGCSE (1)2
2022 Readable vs. Writable Code: A Survey of Intermediate Students' Structure Choices
abstract
Since intermediate CS students can use a variety of control struc- tures, why do their choices often not match experts' Students may not realize what choices expert prefer, find non-expert choices easier to read, or simply forget to write with expert structure. To disentangle these explanations, we surveyed 328 2nd and 3rd se- mester undergraduates, with tasks including writing short func- tions, selecting which structure was most readable or best styled, and comprehension questions. Questions focused on seven control structure topics that were important to instructors (e.g., factoring out repeated code between an if-block and its else). Students frequently wrote with non-expert structure, and, for five topics, at least 1/3 of students (48% - 71%) thought a non-expert struc- ture was more readable than the expert one. However, students often made one choice when writing code, but preferred a different choice when reading it. Additionally, for more complex topics, stu- dents often failed to notice (or understand) differences in execution caused by changes in structure. Together, these results suggest that instruction and practice for choosing control structures should be context-specific, and that assessment focused only on code writing may miss underlying misunderstandings.
Eliane Wiese, Anna N. Rafferty, Jordan Pyper
SIGCSE (1)2
2021 Using Adaptive Experiments to Rapidly Help Students
Angela M. Zavaleta Bernuy, Qi Yin Zheng, Hammad Shaikh, Jacob Nogas, Anna N. Rafferty, Andrew Petersen 0001, Joseph Jay Williams
AIED (2)5
2021 Students' Misunderstanding of the Order of Evaluation in Conjoined Conditions
abstract
Experts often use particular control flow structures to make their code easier to read and modify, such as using the logical operator AND to conjoin conditions rather than nesting separate if statements. Within Boolean expressions, experts take advantage of short-circuit evaluation by ordering their conditions to avoid errors (such as checking that an index is within the bounds of an array before examining the value at that index). How well do students understand these structures? We investigate students' use and understanding of conjoined versus separate conditions within a larger assessment of 125 undergraduate students at the end of their second- and third-semester CS courses (in algorithms & data structures and introductory software engineering). The assessment asked students to: write code where an edge case error could be avoided with short-circuit evaluation, revise their code with nudges towards expert structure, and answer comprehension questions involving code tracing. When writing, students frequently forgot to check for a key edge case. When that case was included, the check was often separated in its own if-statement rather than conjoined with the other conditions. This could indicate a stylistic choice or a belief that the check had to be separated for functionality. Notably, students who included all necessary conditions rarely exhibited the error of ordering them incorrectly. However, with code comprehension, students demonstrated significant misunderstandings about the effects of condition ordering. Students were more accurate on comprehension tasks with nested ifs than conjoined conditions, and this effect was most pronounced when the ordering of the conditions would lead to errors. When conditions were conjoined in a single expression, only 35% of students recognized that checking a value at an index before checking that the index was in bounds would lead to an error. However, 54% of students recognized the problem when the conditions were separated into individual if-statements. This demonstrates a subtlety in code execution that intermediate students may not have mastered and emphasizes the challenges in assessing students' understanding solely via the way they write code.
Eliane Wiese, Anna N. Rafferty, Garrett Moseke
ICPC2
2021 The MOOClet Framework: Unifying Experimentation, Dynamic Improvement, and Personalization in Online Courses
abstract
How can educational platforms be instrumented to accelerate the use of research to improve students' experiences? We show how modular components of any educational interface - e.g. explanations, homework problems, even emails - can be implemented using the novel MOOClet software architecture. Researchers and instructors can use these augmented MOOClet components for: (1) Iterative Cycles of Randomized Experiments that test alternative versions of course content; (2) Data-Driven Improvement using adaptive experiments that rapidly use data to give better versions of content to future students, on the order of days rather than months. A MOOClet supports both manual and automated improvement using reinforcement learning; (3) Personalization by delivering alternative versions as a function of data about a student's characteristics or subgroup, using both expert-authored rules and data mining algorithms. We provide an open-source web service for implementing MOOClets (www.mooclet.org) that has been used with thousands of students. The MOOClet framework provides an ecosystem that transforms online course components into collaborative micro-laboratories, where instructors, experimental researchers, and data mining/machine learning researchers can engage in perpetual cycles of experimentation, improvement, and personalization.
Mohi Reza, Juho Kim 0001, Ananya Bhattacharjee, Anna N. Rafferty, Joseph Jay Williams
L@S4
2020 A rational model of sequential self-assessment
Rachel Jansen, Anna N. Rafferty, Thomas L. Griffiths 0001
CogSci2
2020 Getting too personal(ized): The importance of feature choice in online adaptive algorithms
Zhaobin Li, Luna Yee, Nathaniel Sauerberg, Irene Sakson, Joseph Jay Williams, Anna N. Rafferty
EDM6
2019 Modeling students' fraction arithmetic strategies using inverse planning
Anna N. Rafferty, Rachel Jansen, Thomas L. Griffiths 0001
CogSci1
2019 Balancing Student Success and Inferring Personalized Effects in Dynamic Experiments
Hammad Shaikh, Arghavan Modiri, Joseph Jay Williams, Anna N. Rafferty
EDM4
2019 Replicating novices' struggles with coding style
abstract
Good style makes code easier for others to read and modify. Control flow is one element of style where experts expect particular structure, such as conjoining conditions rather than nesting if statements. Empirical work is necessary to understand why novices use poor style, so they can be taught to use good style. Previous work shows that many students know what control flows experts prefer, but may say that novice-styled code is more readable. Yet, these same students showed similarly high comprehension across both expert-and novice-styled code. We propose a replication of that work that more fully assesses students' code comprehension and code writing. Our replication focuses on students who are earlier in their computer science courses and are less likely to be majors, to determine whether the pattern of results is particular to students who are relatively attuned to style concerns. Our pilot of the proposed replication finds that: students in this new population are less able to identify expert code; expert style may reduce comprehension for some control flows; and writing with good style does not always predict a preference for reading code with good style.
Eliane Wiese, Anna N. Rafferty, Daniel M. Kopta, Jacqulyn M. Anderson
ICPC2
2018 Bandit Assignment for Educational Experiments: Benefits to Students Versus Statistical Power
Anna N. Rafferty, Huiji Ying, Joseph Jay Williams
AIED (2)1
2018 Enhancing Online Problems Through Instructor-Centered Tools for Randomized Experiments
abstract
Digital educational resources could enable the use of randomized experiments to answer pedagogical questions that instructors care about, taking academic research out of the laboratory and into the classroom. We take an instructor-centered approach to designing tools for experimentation that lower the barriers for instructors to conduct experiments. We explore this approach through DynamicProblem, a proof-of-concept system for experimentation on components of digital problems, which provides interfaces for authoring of experiments on explanations, hints, feedback messages, and learning tips. To rapidly turn data from experiments into practical improvements, the system uses an interpretable machine learning algorithm to analyze students' ratings of which conditions are helpful, and present conditions to future students in proportion to the evidence they are higher rated. We evaluated the system by collaboratively deploying experiments in the courses of three mathematics instructors. They reported benefits in reflecting on their pedagogy, and having a new method for improving online problems for future students.
Joseph Jay Williams, Anna N. Rafferty, Dustin Tingley, Andrew M. Ang, Walter S. Lasecki, Juho Kim 0001
CHI2
2018 Modeling the Dunning-Kruger Effect: A Rational Account of Inaccurate Self-Assessment
Rachel Jansen, Anna N. Rafferty, Thomas L. Griffiths 0001
CogSci2
2017 Algebra is not like trivia: Evaluating self-assessment in an online math tutor
Rachel Jansen, Anna N. Rafferty, Thomas L. Griffiths 0001
CogSci2
2017 Eliciting Middle School Students' Ideas About Graphs Supports Their Learning from a Computer Model
Eliane Wiese, Anna N. Rafferty, Marcia C. Linn
CogSci2
2017 MOOClets: A Framework for Dynamic Experimentation and Personalization
abstract
Randomized experiments in online educational environments are ubiquitous as a scientific method for investigating learning and motivation, but too rarely improve educational resources and produce practical benefits for learners. We suggest that software and tools for experimentally comparing resources are designed primarily through the lens of experiments as a scientific methodology, and therefore miss a tremendous opportunity for online experiments to serve as engines for dynamic improvement and personalization. We present the MOOClet requirements specification to guide the implementation of software or tools for experiments to ensure that whenever alternative versions of a resource can be experimentally compared (by randomly assigning versions), the resource can also be dynamically improved (by changing which versions are presented), and personalized (by presenting different versions to different people). The MOOClet specification was used to implement DEXPER, a proof-of-concept web service backend that enables dynamic experimentation and personalization of resources embedded in front-end educational platforms. We describe three use cases of MOOClets for dynamic experimentation and personalization of motivational emails, explanations, and problems.
Joseph Jay Williams, Anna N. Rafferty, Samuel G. Maldonado, Andrew M. Ang, Dustin Tingley, Juho Kim 0001
L@S2
2016 Using Inverse Planning for Personalized Feedback
Anna N. Rafferty, Rachel Jansen, Thomas L. Griffiths 0001
EDM1
2016 AXIS: Generating Explanations at Scale with Learnersourcing and Machine Learning
abstract
While explanations may help people learn by providing information about why an answer is correct, many problems on online platforms lack high-quality explanations. This paper presents AXIS (Adaptive eXplanation Improvement System), a system for obtaining explanations. AXIS asks learners to generate, revise, and evaluate explanations as they solve a problem, and then uses machine learning to dynamically determine which explanation to present to a future learner, based on previous learners' collective input. Results from a case study deployment and a randomized experiment demonstrate that AXIS elicits and identifies explanations that learners find helpful. Providing explanations from AXIS also objectively enhanced learning, when compared to the default practice where learners solved problems and received answers without explanations. The rated quality and learning benefit of AXIS explanations did not differ from explanations generated by an experienced instructor.
Joseph Jay Williams, Juho Kim 0001, Anna N. Rafferty, Samuel G. Maldonado, Krzysztof Z. Gajos, Walter S. Lasecki, Neil T. Heffernan
L@S3
2015 Interpreting Freeform Equation Solving
Anna N. Rafferty, Thomas L. Griffiths 0001
AIED1
2014 The Telephone Game: Exploring Inductive Biases In Naturalistic Language Use
Stephan C. Meylan, Brett Goldstein, Anna N. Rafferty, Thomas L. Griffiths 0001
CogSci3
2014 A Bounded Rationality Account of Wishful Thinking
Rebecca Neumann, Anna N. Rafferty, Thomas L. Griffiths 0001
CogSci2
2014 Tracking Student Understanding of Chemical Reactions in ChemVLab+
Anna N. Rafferty, Jodi L. Davenport, David J. Yaron
CogSci1
2014 Diagnosing Algebra Understanding via Bayesian Inverse Planning
Anna N. Rafferty, Thomas L. Griffiths 0001
EDM1
2013 Estimating Student Knowledge from Paired Interaction Data
Anna N. Rafferty, Jodi L. Davenport, Emma Brunskill
EDM1
2012 Optimally Designing Games for Cognitive Science Research
Anna N. Rafferty, Matei Zaharia, Thomas L. Griffiths 0001
CogSci1
2012 Inferring learners' knowledge from observed actions
Anna N. Rafferty, Michelle M. LaMar, Thomas L. Griffiths 0001
EDM1
2011 Faster Teaching by POMDP Planning
Anna N. Rafferty, Emma Brunskill, Thomas L. Griffiths 0001, Patrick Shafto
AIED1
2008 Finding Contradictions in Text
Marie-Catherine de Marneffe, Anna N. Rafferty, Christopher D. Manning
ACL2
2007 Applying Learning Factors Analysis to Build Stereotypic Student Models
Anna N. Rafferty, Michael Yudelson
AIED1