VLDB 2026 Research / reviewers in the wild / expert
Anna N. Rafferty
dblp:26/6934
· DBLP profile ↗
44ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-8319-5370ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 14 · 5 first-author · 9 since 2021Systems, architecture and hardware · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Likelihood-Based Diagnosis with Generative Models: Confidence-Aware Measurement from Student Writing
Thomas Christie, Matthew Zent, Markus Hauru, Anna N. Rafferty, Simon Woodhead 0002 |
AIED (3) | 4 |
| 2026 | Maximizing the Margin Between Desirable and Undesirable Elements in a Covering Problem
Sophie Boileau, Andrew Hong, David Liben-Nowell, Alistair Pattison, Anna N. Rafferty, Charlie Roslansky |
COCOON | 5 |
| 2025 | An Agentic Framework for Real-Time Pedagogical Plot Generation
Thomas Christie, Anna N. Rafferty, Zack Lee, Ella Cutler, Husni Almoubayyed |
AIED (5) | 2 |
| 2025 | Knowledge of Examples Affects Conditional Reasoning About Math
David W. Braithwaite, Anna N. Rafferty |
CogSci | 2 |
| 2025 | Just Read the Question: Enabling Generalization to New Assessment Items with Text Awareness
Arisha Khan, Nathaniel Li, Tori Shen, Anna N. Rafferty |
EDM | 4 |
| 2025 | Platform-based Adaptive Experimental Research in Education: Lessons Learned from The Digital Learning ChallengeabstractAdaptive Experimentation is one of the most promising approaches to support complex decision-making in learning experience design and delivery. This paper reports on our experience with a real-world, multi-experimental evaluation of an adaptive experimentation platform within the XPRIZE Digital Learning Challenge framework, and summarizes data-driven lessons learned and best practices for Adaptive Experimentation in education. We outline key scenarios of the applicability of platform-supported experiments and reflect on lessons learned from this two-year project, focusing on implications relevant to platform developers, researchers, practitioners, and policy stakeholders to integrate Adaptive Experiments in real-world courses. Ilya Musabirov, Mohi Reza, Haochen Song, Steven Moore, Pan Chen 0005, John C. Stamper, Norman L. Bier, Anna N. Rafferty, Thomas W. Price, Nina Deliu, Audrey Durand, Michael Liut, Joseph Jay Williams |
LAK | 10 |
| 2024 | Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental HealthabstractDigital mental health (DMH) interventions, such as text-message-based lessons and activities, offer immense potential for accessible mental health support. While these interventions can be effective, real-world experimental testing can further enhance their design and impact. Adaptive experimentation, utilizing algorithms like Thompson Sampling for (contextual) multi-armed bandit (MAB) problems, can lead to continuous improvement and personalization. However, it remains unclear when these algorithms can simultaneously increase user experience rewards and facilitate appropriate data collection for social-behavioral scientists to analyze with sufficient statistical confidence. Although a growing body of research addresses the practical and statistical aspects of MAB and other adaptive algorithms, further exploration is needed to assess their impact across diverse real-world contexts. This paper presents a software system developed over two years that allows text-messaging intervention components to be adapted using bandit and other algorithms while collecting data for side-by-side comparison with traditional uniform random non-adaptive experiments. We evaluate the system by deploying a text-message-based DMH intervention to 1100 users, recruited through a large mental health non-profit organization, and share the path forward for deploying this system at scale. This system not only enables applications in mental health but could also serve as a model testbed for adaptive experimentation algorithms in other domains. Jiakai Shi, Ilya Musabirov, Rachel Kornfield, Jonah Meyerhoff, Ananya Bhattacharjee, Chris J. Karr, Theresa Nguyen, David C. Mohr, Anna N. Rafferty, Sofia S. Villar, Nina Deliu, Joseph Jay Williams |
AAAI | 11 |
| 2024 | Uncertainty-preserving deep knowledge tracing with state-space models
Thomas Christie, Carson Cook, Anna N. Rafferty |
EDM | 3 |
| 2024 | Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in ClassroomsabstractSelf-reflection on learning experiences constitutes a fundamental cognitive process, essential for consolidating knowledge and enhancing learning efficacy. However, traditional methods to facilitate reflection often face challenges in personalization, immediacy of feedback, engagement, and scalability. Integration of Large Language Models (LLMs) into the reflection process could mitigate these limitations. In this paper, we conducted two randomized field experiments in undergraduate computer science courses to investigate the potential of LLMs to help students engage in post-lesson reflection. In the first experiment (N=145), students completed a take-home assignment with the support of an LLM assistant; half of these students were then provided access to an LLM designed to facilitate self-reflection. The results indicated that the students assigned to LLM-guided reflection reported somewhat increased self-confidence compared to peers in a no-reflection control and a non-significant trend towards higher scores on a later assessment. Thematic analysis of students' interactions with the LLM showed that the LLM often affirmed the student's understanding, expanded on the student's reflection, and prompted additional reflection; these behaviors suggest ways LLM-interaction might facilitate reflection. In the second experiment (N=112), we evaluated the impact of LLM-guided self-reflection against other scalable reflection methods, such as questionnaire-based activities and review of key lecture slides, after assignment. Our findings suggest that the students in the questionnaire and LLM-based reflection groups performed equally well and better than those who were only exposed to lecture slides, according to their scores on a proctored exam two weeks later on the same subject matter. These results underscore the utility of LLM-guided reflection and questionnaire-based activities in improving learning outcomes. Our work highlights that focusing solely on the accuracy of LLMs can overlook their potential to enhance metacognitive skills through practices such as self-reflection. We discuss the implications of our research for the learning-at-scale community, highlighting the potential of LLMs to enhance learning experiences through personalized, engaging, and scalable reflection practices. Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Huayin Luo, Joseph Jay Williams, Anna N. Rafferty, John C. Stamper, Michael Liut |
L@S | 9 |
| 2024 | Growth in Knowledge of Programming Patterns: A Comparison Study of CS1 vs. CS2 StudentsabstractHow does students' knowledge of code structure improve as they progress through their degree, and where do students struggle? We conducted a comparative study between introductory (CS1) and intermediate CS students (CS2) to explore these questions. Using an online survey with several tasks, including identification of expert patterns, judgment of readable structure, code comprehension, code writing, and editing, we focused on two important code structures: (S1) returning boolean expressions directly and (S2) unique vs. repeated code within if and else. Student performance varied based on structure and task: in both S1 and S2, CS2 students demonstrated higher performance in identifying patterns, judgment of readable structure, and editing. However, evidence of improvement in code writing was only found for S1, and improvement in code comprehension was only found for S2. Therefore, students may need different supports across different code structures. With the exception of comprehension of S1, student performance was far below ceiling, suggesting a need for more support. Sara Nurollahian, Anna N. Rafferty, Noelle Brown, Eliane Wiese |
SIGCSE (1) | 2 |
| 2024 | Playing with Matches: Adopting Gale-Shapley for Managing Student Enrollments Beyond CS2abstractEnrollment in computer science has increased dramatically in recent years, straining capacities and leading to various strategies for managing enrollment. But some strategies increase student competition and may have disproportionate negative impacts on students from underrepresented groups. We believe success in computing education necessitates a more equitable approach to course enrollment. In this experience report, we describe our new enrollment mechanism, "the Match." Building on the Gale--Shapley stable matching algorithm, the Match was designed to encourage a liberal arts approach to course selection and attempt to broaden participation in computing. Drawing on data from three years of use, we find high student participation, with the vast majority of students having their enrollment preferences met. With Match registration, our courses have tended to be a bit more inclusive of younger students. The Match appears not to have disparate negative impacts like those of competitive enrollment, but has increased workload in the Registrar's Office. Overall, we believe the Match has decreased student and faculty angst around registration, and we argue that systems like the Match can help manage enrollment pressures in ways that are consistent with educational values. Anna N. Rafferty, David Liben-Nowell, David R. Musicant, Emy Farley, Allie Lyman, Ann May |
SIGCSE (1) | 1 |
| 2023 | LENS: Predictive Diagnostics for Flexible and Efficient AssessmentsabstractThe utility of assessment systems lies in their capacity to transform observations of student behavior into meaningful inferences about learning, knowledge, and skills. Common practice is to use latent variable models and produce scores on scales. However the simplicity of these psychometric models may filter out potentially valuable information present in student behavior. In particular, scale scores are not optimized to support granular instructional decisions. Machine learning offers promising alternatives, but proposed deep learning architectures are not ideally suited for operational testing conditions involving sparse data and shifting category labels for test questions. Thomas Christie, Hayden Johnson, Carson Cook, Garron Gianopulos, Anna N. Rafferty |
L@S | 5 |
| 2022 | FATED 2022: Fairness, Accountability, and Transparency in Educational Data
Collin F. Lynch, Mirko Marras, Mykola Pechenizkiy, Anna N. Rafferty, Steven Ritter 0001, Vinitra Swamy, Renzhe Yu |
EDM | 4 |
| 2022 | Adversarial bandits for drawing generalizable conclusions in non-adversarial experiments: an empirical study
Zhi-Han Yang, Anna N. Rafferty |
EDM | 3 |
| 2022 | How can Email Interventions Increase Students' Completion of Online Homework? A Case Study Using A/B ComparisonsabstractEmail communication between instructors and students is ubiquitous, and it could be valuable to explore ways of testing out how to make email messages more impactful. This paper explores the design space of using emails to get students to plan and reflect on starting weekly homework earlier. We deployed a series of email reminders using randomized A/B comparisons to test alternative factors in the design of these emails, providing examples of an experimental paradigm and metrics for a broader range of interventions. We also surveyed and interviewed instructors and students to compare their predictions about the effectiveness of the reminders with their actual impact. We present our results on which seemingly obvious predictions about effective emails are not borne out, despite there being evidence for further exploring these interventions, as they can sometimes motivate students to attempt their homework more often. We also present qualitative evidence about student opinions and behaviours after receiving the emails, to guide further interventions. These findings provide insight into how to use randomized A/B comparisons in everyday channels such as emails, to provide empirical evidence to test our beliefs about the effectiveness of alternative design choices. Angela M. Zavaleta Bernuy, Ziwen Han, Hammad Shaikh, Qi Yin Zheng, Lisa-Angelique Lim, Anna N. Rafferty, Andrew Petersen 0001, Joseph Jay Williams |
LAK | 6 |
| 2022 | Student Motivations and Goals for CS1: Themes and VariationsabstractStudents come to CS1 with a wide variety of motivations and goals, which may differ across subpopulations and be indicative of their future engagement with CS. While there is a rich literature relating success in CS1 to specific constructs, such as belonging, goal-orientation, or self-efficacy, less work has examined what motivations and goals students volunteer as most important for their enrollment in CS1. Here, we use qualitative coding to identify themes from students' open-ended descriptions of why they're taking CS1 and what they hope to get out of it, collected across fifteen years. Using quantitative analysis of these coded descriptions, and word-frequency analysis, we identify and name three clusters of students that encompass the majority of students taking CS1: Explorers, Planners, and Utilitarians. We also identify motivations and goals that are more common for particular populations, such as students who have not yet declared a major or students without prior programming experience, as well as factors predicting students' later engagement with CS. This work demonstrates the potential of qualitative coding and computational analyses to enable us to better understand a population of students based on their own words. David Liben-Nowell, Anna N. Rafferty |
SIGCSE (1) | 2 |
| 2022 | Readable vs. Writable Code: A Survey of Intermediate Students' Structure ChoicesabstractSince intermediate CS students can use a variety of control struc- tures, why do their choices often not match experts' Students may not realize what choices expert prefer, find non-expert choices easier to read, or simply forget to write with expert structure. To disentangle these explanations, we surveyed 328 2nd and 3rd se- mester undergraduates, with tasks including writing short func- tions, selecting which structure was most readable or best styled, and comprehension questions. Questions focused on seven control structure topics that were important to instructors (e.g., factoring out repeated code between an if-block and its else). Students frequently wrote with non-expert structure, and, for five topics, at least 1/3 of students (48% - 71%) thought a non-expert struc- ture was more readable than the expert one. However, students often made one choice when writing code, but preferred a different choice when reading it. Additionally, for more complex topics, stu- dents often failed to notice (or understand) differences in execution caused by changes in structure. Together, these results suggest that instruction and practice for choosing control structures should be context-specific, and that assessment focused only on code writing may miss underlying misunderstandings. Eliane Wiese, Anna N. Rafferty, Jordan Pyper |
SIGCSE (1) | 2 |
| 2021 | Using Adaptive Experiments to Rapidly Help Students
Angela M. Zavaleta Bernuy, Qi Yin Zheng, Hammad Shaikh, Jacob Nogas, Anna N. Rafferty, Andrew Petersen 0001, Joseph Jay Williams |
AIED (2) | 5 |
| 2021 | Students' Misunderstanding of the Order of Evaluation in Conjoined ConditionsabstractExperts often use particular control flow structures to make their code easier to read and modify, such as using the logical operator AND to conjoin conditions rather than nesting separate if statements. Within Boolean expressions, experts take advantage of short-circuit evaluation by ordering their conditions to avoid errors (such as checking that an index is within the bounds of an array before examining the value at that index). How well do students understand these structures? We investigate students' use and understanding of conjoined versus separate conditions within a larger assessment of 125 undergraduate students at the end of their second- and third-semester CS courses (in algorithms & data structures and introductory software engineering). The assessment asked students to: write code where an edge case error could be avoided with short-circuit evaluation, revise their code with nudges towards expert structure, and answer comprehension questions involving code tracing. When writing, students frequently forgot to check for a key edge case. When that case was included, the check was often separated in its own if-statement rather than conjoined with the other conditions. This could indicate a stylistic choice or a belief that the check had to be separated for functionality. Notably, students who included all necessary conditions rarely exhibited the error of ordering them incorrectly. However, with code comprehension, students demonstrated significant misunderstandings about the effects of condition ordering. Students were more accurate on comprehension tasks with nested ifs than conjoined conditions, and this effect was most pronounced when the ordering of the conditions would lead to errors. When conditions were conjoined in a single expression, only 35% of students recognized that checking a value at an index before checking that the index was in bounds would lead to an error. However, 54% of students recognized the problem when the conditions were separated into individual if-statements. This demonstrates a subtlety in code execution that intermediate students may not have mastered and emphasizes the challenges in assessing students' understanding solely via the way they write code. Eliane Wiese, Anna N. Rafferty, Garrett Moseke |
ICPC | 2 |
| 2021 | The MOOClet Framework: Unifying Experimentation, Dynamic Improvement, and Personalization in Online CoursesabstractHow can educational platforms be instrumented to accelerate the use of research to improve students' experiences? We show how modular components of any educational interface - e.g. explanations, homework problems, even emails - can be implemented using the novel MOOClet software architecture. Researchers and instructors can use these augmented MOOClet components for: (1) Iterative Cycles of Randomized Experiments that test alternative versions of course content; (2) Data-Driven Improvement using adaptive experiments that rapidly use data to give better versions of content to future students, on the order of days rather than months. A MOOClet supports both manual and automated improvement using reinforcement learning; (3) Personalization by delivering alternative versions as a function of data about a student's characteristics or subgroup, using both expert-authored rules and data mining algorithms. We provide an open-source web service for implementing MOOClets (www.mooclet.org) that has been used with thousands of students. The MOOClet framework provides an ecosystem that transforms online course components into collaborative micro-laboratories, where instructors, experimental researchers, and data mining/machine learning researchers can engage in perpetual cycles of experimentation, improvement, and personalization. Mohi Reza, Juho Kim 0001, Ananya Bhattacharjee, Anna N. Rafferty, Joseph Jay Williams |
L@S | 4 |
| 2020 | A rational model of sequential self-assessment
Rachel Jansen, Anna N. Rafferty, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2020 | Getting too personal(ized): The importance of feature choice in online adaptive algorithms
Zhaobin Li, Luna Yee, Nathaniel Sauerberg, Irene Sakson, Joseph Jay Williams, Anna N. Rafferty |
EDM | 6 |
| 2019 | Modeling students' fraction arithmetic strategies using inverse planning
Anna N. Rafferty, Rachel Jansen, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2019 | Balancing Student Success and Inferring Personalized Effects in Dynamic Experiments
Hammad Shaikh, Arghavan Modiri, Joseph Jay Williams, Anna N. Rafferty |
EDM | 4 |
| 2019 | Replicating novices' struggles with coding styleabstractGood style makes code easier for others to read and modify. Control flow is one element of style where experts expect particular structure, such as conjoining conditions rather than nesting if statements. Empirical work is necessary to understand why novices use poor style, so they can be taught to use good style. Previous work shows that many students know what control flows experts prefer, but may say that novice-styled code is more readable. Yet, these same students showed similarly high comprehension across both expert-and novice-styled code. We propose a replication of that work that more fully assesses students' code comprehension and code writing. Our replication focuses on students who are earlier in their computer science courses and are less likely to be majors, to determine whether the pattern of results is particular to students who are relatively attuned to style concerns. Our pilot of the proposed replication finds that: students in this new population are less able to identify expert code; expert style may reduce comprehension for some control flows; and writing with good style does not always predict a preference for reading code with good style. Eliane Wiese, Anna N. Rafferty, Daniel M. Kopta, Jacqulyn M. Anderson |
ICPC | 2 |
| 2018 | Bandit Assignment for Educational Experiments: Benefits to Students Versus Statistical Power
Anna N. Rafferty, Huiji Ying, Joseph Jay Williams |
AIED (2) | 1 |
| 2018 | Enhancing Online Problems Through Instructor-Centered Tools for Randomized ExperimentsabstractDigital educational resources could enable the use of randomized experiments to answer pedagogical questions that instructors care about, taking academic research out of the laboratory and into the classroom. We take an instructor-centered approach to designing tools for experimentation that lower the barriers for instructors to conduct experiments. We explore this approach through DynamicProblem, a proof-of-concept system for experimentation on components of digital problems, which provides interfaces for authoring of experiments on explanations, hints, feedback messages, and learning tips. To rapidly turn data from experiments into practical improvements, the system uses an interpretable machine learning algorithm to analyze students' ratings of which conditions are helpful, and present conditions to future students in proportion to the evidence they are higher rated. We evaluated the system by collaboratively deploying experiments in the courses of three mathematics instructors. They reported benefits in reflecting on their pedagogy, and having a new method for improving online problems for future students. Joseph Jay Williams, Anna N. Rafferty, Dustin Tingley, Andrew M. Ang, Walter S. Lasecki, Juho Kim 0001 |
CHI | 2 |
| 2018 | Modeling the Dunning-Kruger Effect: A Rational Account of Inaccurate Self-Assessment
Rachel Jansen, Anna N. Rafferty, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2017 | Algebra is not like trivia: Evaluating self-assessment in an online math tutor
Rachel Jansen, Anna N. Rafferty, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2017 | Eliciting Middle School Students' Ideas About Graphs Supports Their Learning from a Computer Model
Eliane Wiese, Anna N. Rafferty, Marcia C. Linn |
CogSci | 2 |
| 2017 | MOOClets: A Framework for Dynamic Experimentation and PersonalizationabstractRandomized experiments in online educational environments are ubiquitous as a scientific method for investigating learning and motivation, but too rarely improve educational resources and produce practical benefits for learners. We suggest that software and tools for experimentally comparing resources are designed primarily through the lens of experiments as a scientific methodology, and therefore miss a tremendous opportunity for online experiments to serve as engines for dynamic improvement and personalization. We present the MOOClet requirements specification to guide the implementation of software or tools for experiments to ensure that whenever alternative versions of a resource can be experimentally compared (by randomly assigning versions), the resource can also be dynamically improved (by changing which versions are presented), and personalized (by presenting different versions to different people). The MOOClet specification was used to implement DEXPER, a proof-of-concept web service backend that enables dynamic experimentation and personalization of resources embedded in front-end educational platforms. We describe three use cases of MOOClets for dynamic experimentation and personalization of motivational emails, explanations, and problems. Joseph Jay Williams, Anna N. Rafferty, Samuel G. Maldonado, Andrew M. Ang, Dustin Tingley, Juho Kim 0001 |
L@S | 2 |
| 2016 | Using Inverse Planning for Personalized Feedback
Anna N. Rafferty, Rachel Jansen, Thomas L. Griffiths 0001 |
EDM | 1 |
| 2016 | AXIS: Generating Explanations at Scale with Learnersourcing and Machine LearningabstractWhile explanations may help people learn by providing information about why an answer is correct, many problems on online platforms lack high-quality explanations. This paper presents AXIS (Adaptive eXplanation Improvement System), a system for obtaining explanations. AXIS asks learners to generate, revise, and evaluate explanations as they solve a problem, and then uses machine learning to dynamically determine which explanation to present to a future learner, based on previous learners' collective input. Results from a case study deployment and a randomized experiment demonstrate that AXIS elicits and identifies explanations that learners find helpful. Providing explanations from AXIS also objectively enhanced learning, when compared to the default practice where learners solved problems and received answers without explanations. The rated quality and learning benefit of AXIS explanations did not differ from explanations generated by an experienced instructor. Joseph Jay Williams, Juho Kim 0001, Anna N. Rafferty, Samuel G. Maldonado, Krzysztof Z. Gajos, Walter S. Lasecki, Neil T. Heffernan |
L@S | 3 |
| 2015 | Interpreting Freeform Equation Solving
Anna N. Rafferty, Thomas L. Griffiths 0001 |
AIED | 1 |
| 2014 | The Telephone Game: Exploring Inductive Biases In Naturalistic Language Use
Stephan C. Meylan, Brett Goldstein, Anna N. Rafferty, Thomas L. Griffiths 0001 |
CogSci | 3 |
| 2014 | A Bounded Rationality Account of Wishful Thinking
Rebecca Neumann, Anna N. Rafferty, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2014 | Tracking Student Understanding of Chemical Reactions in ChemVLab+
Anna N. Rafferty, Jodi L. Davenport, David J. Yaron |
CogSci | 1 |
| 2014 | Diagnosing Algebra Understanding via Bayesian Inverse Planning
Anna N. Rafferty, Thomas L. Griffiths 0001 |
EDM | 1 |
| 2013 | Estimating Student Knowledge from Paired Interaction Data
Anna N. Rafferty, Jodi L. Davenport, Emma Brunskill |
EDM | 1 |
| 2012 | Optimally Designing Games for Cognitive Science Research
Anna N. Rafferty, Matei Zaharia, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2012 | Inferring learners' knowledge from observed actions
Anna N. Rafferty, Michelle M. LaMar, Thomas L. Griffiths 0001 |
EDM | 1 |
| 2011 | Faster Teaching by POMDP Planning
Anna N. Rafferty, Emma Brunskill, Thomas L. Griffiths 0001, Patrick Shafto |
AIED | 1 |
| 2008 | Finding Contradictions in Text
Marie-Catherine de Marneffe, Anna N. Rafferty, Christopher D. Manning |
ACL | 2 |
| 2007 | Applying Learning Factors Analysis to Build Stereotypic Student Models
Anna N. Rafferty, Michael Yudelson |
AIED | 1 |