VLDB 2026 Research / reviewers in the wild / expert
Joseph Jay Williams
dblp:132/4086
· DBLP profile ↗
82ranked-venue papers
12as first author
45since 2021 · last 2025
0000-0002-9122-5242ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 45 · 3 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 9 first-author · 14 since 2021Artificial intelligence and machine learning · 30 · 8 first-author · 12 since 2021Systems, architecture and hardware · 16 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perfectly to a Tee: Understanding User Perceptions of Personalized LLM-Enhanced Narrative InterventionsabstractStories about overcoming personal struggles can effectively illustrate the application of psychological theories in real life, yet they may fail to resonate with individuals' experiences. In this work, we employ large language models (LLMs) to create tailored narratives that acknowledge and address unique challenging thoughts and situations faced by individuals. Our study, involving 346 young adults across two settings, demonstrates that personalized LLM-enhanced stories were perceived to be better than human-written ones in conveying key takeaways, promoting reflection, and reducing belief in negative thoughts. These stories were not only seen as more relatable but also similarly authentic to human-written ones, highlighting the potential of LLMs in helping young adults manage their struggles. The findings of this work provide crucial design considerations for future narrative-based digital mental health interventions, such as the need to maintain relatability without veering into implausibility and refining the wording and tone of AI-enhanced content. Ananya Bhattacharjee, Sarah Yi Xu, Pranav Rao, Yuchen Zeng 0001, Jonah Meyerhoff, Syed Ishtiaque Ahmed, David C. Mohr, Michael Liut, Alexander Mariakakis, Rachel Kornfield, Joseph Jay Williams |
Conference on Designing Interactive Systems | 11 |
| 2025 | Platform-based Adaptive Experimental Research in Education: Lessons Learned from The Digital Learning ChallengeabstractAdaptive Experimentation is one of the most promising approaches to support complex decision-making in learning experience design and delivery. This paper reports on our experience with a real-world, multi-experimental evaluation of an adaptive experimentation platform within the XPRIZE Digital Learning Challenge framework, and summarizes data-driven lessons learned and best practices for Adaptive Experimentation in education. We outline key scenarios of the applicability of platform-supported experiments and reflect on lessons learned from this two-year project, focusing on implications relevant to platform developers, researchers, practitioners, and policy stakeholders to integrate Adaptive Experiments in real-world courses. Ilya Musabirov, Mohi Reza, Haochen Song, Steven Moore, Pan Chen 0005, John C. Stamper, Norman L. Bier, Anna N. Rafferty, Thomas W. Price, Nina Deliu, Audrey Durand, Michael Liut, Joseph Jay Williams |
LAK | 15 |
| 2025 | Sixth Annual Workshop on A/B Testing and Platform-Enabled Learning Engineering (PELE)abstractLearning engineering applies data and learning science principles to better understand outcomes and support improvement research. One important approach is A/B testing-common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE), and the International Consortium for Innovation and Collaboration in Learning Engineering (IEEE ICICLE). Several systems supporting A/B testing in educational applications have arisen recently, including UpGrade, E-TRIALS, and Terracotta. A/B testing can help improve educational platforms, yet there are challenging issues unique to conducting such work in these contexts. In response, a number of digital learning platforms have opened their systems to learning-improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore how A/B testing is conducted in educational contexts, how digital learning platforms are accelerating education research, and how empirical approaches can be used to drive powerful gains in student learning. It will also discuss opportunities for funding to conduct platform-enabled learning engineering. April Murphy, Stephen Fancsali, Steven Ritter 0001, Neil T. Heffernan, Debshila Basu Mallick, Jeremy Roschelle, Danielle S. McNamara, Joseph Jay Williams, John C. Stamper, Norman L. Bier, Jeffrey C. Carver |
L@S | 8 |
| 2025 | Investigating the Role of Situational Disruptors in Engagement with Digital Mental Health ToolsabstractChallenges in engagement with digital mental health (DMH) tools are commonly addressed through technical enhancements and algorithmic interventions. This paper shifts the focus towards the role of users' broader social context as a significant factor in engagement. Through an eight-week text messaging program aimed at enhancing psychological wellbeing, we recruited 20 participants to help us identify situational engagement disruptors (SEDs), including personal responsibilities, professional obligations, and unexpected health issues. In follow-up design workshops with 25 participants, we explored potential solutions that address such SEDs: prioritizing self-care through structured goal-setting, alternative framings for disengagement, and utilization of external resources. Our findings challenge conventional perspectives on engagement and offer actionable design implications for future DMH tools. Ananya Bhattacharjee, Joseph Jay Williams, Miranda L. Beltzer, Jonah Meyerhoff, Haochen Song, David C. Mohr, Alexander Mariakakis, Rachel Kornfield |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2025 | Large Language Model Agents for Improving Engagement with Behavior Change Interventions: Application to Digital MindfulnessabstractAlthough engagement in self-directed wellness exercises typically declines over time, integrating social support such as coaching can sustain it. However, traditional forms of support are often inaccessible due to the high costs and complex coordination. Large Language Models (LLMs) show promise in providing human-like dialogues that could emulate social support. Yet, in-depth, in situ investigations of LLMs to support behavior change remain underexplored. We conducted two randomized experiments to assess the impact of LLM agents on user engagement with mindfulness exercises. First, a single-session study, involved 502 crowdworkers; second, a three-week study, included 54 participants. We explored two types of LLM agents: one providing information and another facilitating self-reflection. Both agents enhanced users' intentions to practice mindfulness. However, only the information-providing LLM agent, featuring a friendly persona, significantly improved engagement with the exercises. Our findings suggest that specific LLM agents may bridge the social support gap in digital health interventions. Suhyeon Yoo, Angela M. Zavaleta Bernuy, Jiakai Shi, Huayin Luo, Joseph Jay Williams, Anastasia Kuzminykh, Ashton Anderson, Rachel Kornfield |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2025 | Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing ProcessabstractAs generative AI tools like ChatGPT become integral to everyday writing, critical questions arise about how to preserve writers' sense of agency and ownership when using these tools. Yet, a systematic understanding of how AI assistance affects different aspects of the writing process-and how this shapes writers' agency-remains underexplored. To address this gap, we conducted a systematic review of 109 HCI papers using the PRISMA approach. From this literature, we identify four overarching design strategies for AI writing support- structured guidance, guided exploration, active co-writing , and critical feedback -mapped across the four key cognitive processes in writing: planning, translating, reviewing , and monitoring . We complement this analysis with interviews of 15 writers across diverse domains. Our findings reveal that writers' desired levels of AI intervention vary across the writing process: content-focused writers (e.g., academics) prioritize ownership during planning, while form-focused writers (e.g., creatives) value control over translating and reviewing. Writers' preferences are also shaped by contextual goals, values, and notions of originality and authorship. By examining when ownership matters, what writers want to own, and how AI interactions shape agency, we surface both alignment and gaps between research and user needs. Our findings offer actionable design guidance for developing human-centered writing tools for co-writing with AI, on human terms. Mohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry, Joseph Jay Williams, Anastasia Kuzminykh |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Does the Medium Matter? An Exploration of Voice-Interaction for Self-ExplanationsabstractThis research evaluates voice-based self-explanations as a pedagogical tool in preparation for lectures, assesses user preferences between voice and text, and derives design insights. We report two studies: Study 1, a quasi-experimental field study, with 247 participants divided into voice-based (N = 83), text-based (N = 81), and choice (N = 83) conditions. Study 2 uses semi-structured interviews (N = 16) to explore perceptions of the interaction paradigms in-depth. Results from the first study revealed a general preference for text, though voice users produced longer responses and more topic-related keywords. Over time, the preference for voice increased among students, from 10% to 46%, when given a choice. Study 2 suggested that factors like social presence contribute to hesitance toward voice-based explanations, with a cognitive load, self-confidence, and performance anxiety also influencing medium preferences. Our findings highlight design recommendations and demonstrate the potential of voice-based self-explanations in educational settings, indicating that mixed interfaces might better meet diverse needs. Angela M. Zavaleta Bernuy, Naaz Sibia, Pan Chen 0005, Jessica Jia-Ni Xu, Elexandra Tran, Runlong Ye 0002, Viktoria Pammer-Schindler, Andrew Petersen 0001, Joseph Jay Williams, Michael Liut |
Conference on Designing Interactive Systems | 9 |
| 2024 | Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental HealthabstractDigital mental health (DMH) interventions, such as text-message-based lessons and activities, offer immense potential for accessible mental health support. While these interventions can be effective, real-world experimental testing can further enhance their design and impact. Adaptive experimentation, utilizing algorithms like Thompson Sampling for (contextual) multi-armed bandit (MAB) problems, can lead to continuous improvement and personalization. However, it remains unclear when these algorithms can simultaneously increase user experience rewards and facilitate appropriate data collection for social-behavioral scientists to analyze with sufficient statistical confidence. Although a growing body of research addresses the practical and statistical aspects of MAB and other adaptive algorithms, further exploration is needed to assess their impact across diverse real-world contexts. This paper presents a software system developed over two years that allows text-messaging intervention components to be adapted using bandit and other algorithms while collecting data for side-by-side comparison with traditional uniform random non-adaptive experiments. We evaluate the system by deploying a text-message-based DMH intervention to 1100 users, recruited through a large mental health non-profit organization, and share the path forward for deploying this system at scale. This system not only enables applications in mental health but could also serve as a model testbed for adaptive experimentation algorithms in other domains. Jiakai Shi, Ilya Musabirov, Rachel Kornfield, Jonah Meyerhoff, Ananya Bhattacharjee, Chris J. Karr, Theresa Nguyen, David C. Mohr, Anna N. Rafferty, Sofia S. Villar, Nina Deliu, Joseph Jay Williams |
AAAI | 14 |
| 2024 | Understanding the Role of Large Language Models in Personalizing and Scaffolding Strategies to Combat Academic ProcrastinationabstractTraditional interventions for academic procrastination often fail to capture the nuanced, individual-specific factors that underlie them. Large language models (LLMs) hold immense potential for addressing this gap by permitting open-ended inputs, including the ability to customize interventions to individuals' unique needs. However, user expectations and potential limitations of LLMs in this context remain underexplored. To address this, we conducted interviews and focus group discussions with 15 university students and 6 experts, during which a technology probe for generating personalized advice for managing procrastination was presented. Our results highlight the necessity for LLMs to provide structured, deadline-oriented steps and enhanced user support mechanisms. Additionally, our results surface the need for an adaptive approach to questioning based on factors like busyness. These findings offer crucial design implications for the development of LLM-based tools for managing procrastination while cautioning the use of LLMs for therapeutic guidance. Ananya Bhattacharjee, Yuchen Zeng 0001, Sarah Yi Xu, Dana Kulzhabayeva, Minyi Ma, Rachel Kornfield, Syed Ishtiaque Ahmed, Alexander Mariakakis, Mary Czerwinski, Anastasia Kuzminykh, Michael Liut, Joseph Jay Williams |
CHI | 12 |
| 2024 | ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Language ModelsabstractExploring alternative ideas by rewriting text is integral to the writing process. State-of-the-art Large Language Models (LLMs) can simplify writing variation generation. However, current interfaces pose challenges for simultaneous consideration of multiple variations: creating new variations without overwriting text can be difficult, and pasting them sequentially can clutter documents, increasing workload and disrupting writers’ flow. To tackle this, we present ABScribe, an interface that supports rapid, yet visually structured, exploration and organization of writing variations in human-AI co-writing tasks. With ABScribe, users can swiftly modify variations using LLM prompts, which are auto-converted into reusable buttons. Variations are stored adjacently within text fields for rapid in-place comparisons using mouse-over interactions on a popup toolbar. Our user study with 12 writers shows that ABScribe significantly reduces task workload (d = 1.20, p < 0.001), enhances user perceptions of the revision process (d = 2.41, p < 0.001) compared to a popular baseline workflow, and provides insights into how writers explore variations using LLMs. Mohi Reza, Nathan Laundry, Ilya Musabirov, Peter Dushniku, Zhi Yuan "Michael" Yu, Kashish Mittal, Tovi Grossman, Michael Liut, Anastasia Kuzminykh, Joseph Jay Williams |
CHI | 10 |
| 2024 | Dynamics of Causal Attribution
Dana Kulzhabayeva, Joseph Jay Williams, David Danks |
CogSci | 2 |
| 2024 | Fifth Annual Workshop on A/B Testing and Platform-Enabled Learning ResearchabstractLearning engineering adds tools and processes to learning platforms to support improvement research. One kind of tool is A/B testing-common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE), and the International Consortium for Innovation and Collaboration in Learning Engineering (IEEE ICICLE). Recently, several A/B testing systems have arisen that focus on conducting research in educational environments, including UpGrade, Terracotta, and E-TRIALS. A/B testing can help improve educational platforms, yet there are challenging issues unique to conducting such work in these contexts. In response, a number of digital learning platforms have opened their systems to learning-improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore challenges of A/B testing in educational contexts, how learning platforms are accelerating education research, and how empirical approaches can be used to drive powerful gains in student learning. It will also discuss opportunities for funding to conduct platform-enabled learning research. Steven Ritter 0001, Stephen Fancsali, April Murphy, Neil T. Heffernan, Benjamin Motz 0002, Debshila Basu Mallick, Jeremy Roschelle, Danielle S. McNamara, Joseph Jay Williams |
L@S | 9 |
| 2024 | Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in ClassroomsabstractSelf-reflection on learning experiences constitutes a fundamental cognitive process, essential for consolidating knowledge and enhancing learning efficacy. However, traditional methods to facilitate reflection often face challenges in personalization, immediacy of feedback, engagement, and scalability. Integration of Large Language Models (LLMs) into the reflection process could mitigate these limitations. In this paper, we conducted two randomized field experiments in undergraduate computer science courses to investigate the potential of LLMs to help students engage in post-lesson reflection. In the first experiment (N=145), students completed a take-home assignment with the support of an LLM assistant; half of these students were then provided access to an LLM designed to facilitate self-reflection. The results indicated that the students assigned to LLM-guided reflection reported somewhat increased self-confidence compared to peers in a no-reflection control and a non-significant trend towards higher scores on a later assessment. Thematic analysis of students' interactions with the LLM showed that the LLM often affirmed the student's understanding, expanded on the student's reflection, and prompted additional reflection; these behaviors suggest ways LLM-interaction might facilitate reflection. In the second experiment (N=112), we evaluated the impact of LLM-guided self-reflection against other scalable reflection methods, such as questionnaire-based activities and review of key lecture slides, after assignment. Our findings suggest that the students in the questionnaire and LLM-based reflection groups performed equally well and better than those who were only exposed to lecture slides, according to their scores on a proctored exam two weeks later on the same subject matter. These results underscore the utility of LLM-guided reflection and questionnaire-based activities in improving learning outcomes. Our work highlights that focusing solely on the accuracy of LLMs can overlook their potential to enhance metacognitive skills through practices such as self-reflection. We discuss the implications of our research for the learning-at-scale community, highlighting the potential of LLMs to enhance learning experiences through personalized, engaging, and scalable reflection practices. Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Huayin Luo, Joseph Jay Williams, Anna N. Rafferty, John C. Stamper, Michael Liut |
L@S | 8 |
| 2024 | Student Interaction with Instructor Emails in Introductory and Upper-Year Computing CoursesabstractIn computing courses, instructor involvement and social comfort are vital for resilience and belonging. We examine engagement with instructor emails aimed at strengthening the connection with students. We sent weekly emails from instructors to first- and upper-year computing students. These emails included reminders for the assignments due each week. Half of the students received reminders embedded in an informal message that contained approachable wording and relevant current course events, while the rest received a list of precise deadlines. This text had no emotional engagement from the instructor. We collected and analyzed email access and link click rates, along with student survey responses about email preferences and engagement. We found that first-year students had lower email access and link click rates than upper-year students. While we did not find differences in first-year engagement based on the type of email, upper-year students appeared to be more engaged when receiving the intentionally informal version of the email. Understanding the message preferences of computing students can enhance instructor messaging and improve engagement. Strategies should be explored to boost first-year student engagement, while the higher engagement among upper-year students underscores the importance of instructor support in advanced courses. Angela M. Zavaleta Bernuy, Runlong Ye 0002, Naaz Sibia, Rohita Nalluri, Joseph Jay Williams, Andrew Petersen 0001, Bogdan Simion, Michael Liut |
SIGCSE (1) | 5 |
| 2024 | Examining Intention to Major in Computer Science: Perceived Potential and ChallengesabstractThis study explores links between attributes of computing students, such as prior programming experience (PE) and gender, with expectations for success and the perception of challenges. Using Expectancy-Value Theory (EVT), we investigate their major intentions and the impact of these factors post-CS1. Data was gathered using surveys at the beginning and end of an introductory programming course, focusing on demographics, expectations of success, and perceptions of challenges. Application status for the computing major was also recorded. Our results revealed that men and students with PE generally perceived greater potential for success and reported facing fewer challenges. In contrast, women and students without PE more often indicated concerns about intellectual ability and perceived challenges less positively. Notably, while gender appears in the preceding results, an intersectional analysis indicates that PE is the central factor. PE is also linked to persistence in the field of computing. Our results further highlight the importance of providing students with opportunities to develop experience, as it can help shape their expectations, perceived challenges, and retention in computing. Naaz Sibia, Giang Bui, Bingcheng Wang, Yinyue Tan, Angela M. Zavaleta Bernuy, Christina Bauer, Joseph Jay Williams, Michael Liut, Andrew Petersen 0001 |
SIGCSE (1) | 7 |
| 2024 | "Actually I Can Count My Blessings": User-Centered Design of an Application to Promote Gratitude Among Young AdultsabstractRegular practice of gratitude has the potential to enhance psychological wellbeing and foster stronger social connections among young adults. However, there is a lack of research investigating user needs and expectations regarding gratitude-promoting applications. To address this gap, we employed a user-centered design approach to develop a mobile application that facilitates gratitude practice. Our formative study involved 20 participants who utilized an existing application, providing insights into their preferences for organizing expressions of gratitude and the significance of prompts for reflection and mood labeling after working hours. Building on these findings, we conducted a deployment study with 26 participants using our custom-designed application, which confirmed the positive impact of structured options to guide gratitude practice and highlighted the advantages of passive engagement with the application during busy periods. Our study contributes to the field by identifying key design considerations for promoting gratitude among young adults. Ananya Bhattacharjee, Zichen Gong, Bingcheng Wang, Timothy James Luckcock, Emma Watson, Elena Allica Abellan, Leslie Gutman, Anne Hsu, Joseph Jay Williams |
Proc. ACM Hum. Comput. Interact. | 9 |
| 2024 | Guiding Students in Using LLMs in Supported Learning Environments: Effects on Interaction Dynamics, Learner Performance, Confidence, and TrustabstractPersonalized chatbot-based teaching assistants can be crucial in addressing increasing classroom sizes, especially where direct teacher presence is limited. Large language models (LLMs) offer a promising avenue, with increasing research exploring their educational utility. However, the challenge lies not only in establishing the efficacy of LLMs but also in discerning the nuances of interaction between learners and these models, which impact learners' engagement and results. We conducted a formative study in an undergraduate computer science classroom (N=145) and a controlled experiment on Prolific (N=356) to explore the impact of four pedagogically informed guidance strategies on the learners' performance, confidence and trust in LLMs. Direct LLM answers marginally improved performance, while refining student solutions fostered trust. Structured guidance reduced random queries as well as instances of students copy-pasting assignment questions to the LLM. Our work highlights the role that teachers can play in shaping LLM-supported learning environments. Ilya Musabirov, Mohi Reza, Jiakai Shi, Joseph Jay Williams, Anastasia Kuzminykh, Michael Liut |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2023 | Investigating the Role of Context in the Delivery of Text Messages for Supporting Psychological WellbeingabstractWithout a nuanced understanding of users' perspectives and contexts, text messaging tools for supporting psychological wellbeing risk delivering interventions that are mismatched to users' dynamic needs. We investigated the contextual factors that influence young adults' day-to-day experiences when interacting with such tools. Through interviews and focus group discussions with 36 participants, we identified that people's daily schedules and affective states were dominant factors that shape their messaging preferences. We developed two messaging dialogues centered around these factors, which we deployed to 42 participants to test and extend our initial understanding of users' needs. Across both studies, participants provided diverse opinions of how they could be best supported by messages, particularly around when to engage users in more passive versus active ways. They also proposed ways of adjusting message length and content during periods of low mood. Our findings provide design implications and opportunities for context-aware mental health management systems. Ananya Bhattacharjee, Joseph Jay Williams, Jonah Meyerhoff, Alexander Mariakakis, Rachel Kornfield |
CHI | 2 |
| 2023 | Exam Eustress: Designing Brief Online Interventions for Helping Students Identify Positive Aspects of StressabstractStress reappraisal interventions try to shift students’ negative perceptions towards eustress, stress that can be beneficial, and help them perform better. However, it is less clear how to present them to users as online interventions that are brief, voluntary, and scale well in real-world contexts. We explore the design of online exam eustress interventions by generating six design factors (D1-6) that reinforce a core reappraisal message (D0), and evaluate them through: (i) user interviews (N = 20) revealing six findings (F1-6) on the importance of elaboration, layout, modality, and source of intervention content; (ii) a field experiment (N = 1283) showing a significant positive effect on exam scores (p = 0.003). Subgroup analyses indicate a significant effect for first-year but not for upper-year students, and no detectable gender differences. Our work offers insight into how students interact with online mindset interventions and design considerations for incorporating them into large courses. Mohi Reza, Angela M. Zavaleta Bernuy, Emmy Liu, Zhongyuan Liang, Calista K. Barber, Joseph Jay Williams |
CHI | 7 |
| 2023 | VoiceEx: Voice Submission System for Interventions in EducationabstractGenerating self-explanations has been identified as a successful strategy in helping learners engage with course content and organize what they learn in a structured format. While typing an explanation may allow more structure and formality, explaining by voice can be more natural and help free cognitive resources to focus on learning goals and understanding concepts. As we investigated the effects and students' perceptions of using voice or text to self-explain new course concepts, we failed to find a tool that would meet our needs. We present our work in designing and developing VoiceEx, a submission courseware that allows text and voice input to collect data in both mediums. VoiceEx was created to support a self-explanations intervention for computer science students; however, given its features and the advantages of being able to collect spoken responses, it can be used in a variety of environments. Future refinement of this tool includes artificial intelligence features to better guide students' submissions. Angela M. Zavaleta Bernuy, Naaz Sibia, Pan Chen 0005, Chloe Huang, Andrew Petersen 0001, Joseph Jay Williams, Michael Liut |
ITiCSE (2) | 6 |
| 2023 | Self-Explanation Modality: Effects on Student Performance?abstractIn this poster, we present a pilot study investigating the impact of the medium used for self-explanation on students' performance outcomes in a databases course. We did not see notable differences in student performance based on the medium they used for self-explanation. We also note that while most students prefer using text to submit self-explanations, their preferences may differ when they use voice. Angela M. Zavaleta Bernuy, Jessica Jia-Ni Xu, Naaz Sibia, Joseph Jay Williams, Andrew Petersen 0001, Michael Liut |
ITiCSE (2) | 4 |
| 2023 | Student Usage of Q&A Forums: Signs of Discomfort?abstractQ&A forums are widely used in large classes to provide scalable support. In addition to offering students a space to ask questions, these forums aim to create a community and promote engagement. Prior literature suggests that the way students participate in Q&A forums varies and that most students do not actively post questions or engage in discussions. Students may display different participation behaviours depending on their comfort levels in the class. This paper investigates students' use of a Q&A forum in a CS1 course. We also analyze student opinions about the forum to explain the observed behaviour, focusing on students' lack of visible participation (lurking, anonymity, private posting). We analyzed forum data collected in a CS1 course across two consecutive years and invited students to complete a survey about perspectives on their forum usage. Despite a small cohort of highly engaged students, we confirmed that most students do not actively read or post on the forum. We discuss students' reasons for the low level of engagement and barriers to participating visibly. Common reasons include fearing a lack of knowledge and repercussions from being visible to the student community. Naaz Sibia, Angela M. Zavaleta Bernuy, Joseph Jay Williams, Michael Liut, Andrew Petersen 0001 |
ITiCSE (1) | 3 |
| 2023 | Fourth Annual Workshop on A/B Testing and Platform-Enabled Learning Research
Steven Ritter 0001, Neil T. Heffernan, Joseph Jay Williams, Derek Lomas, Klinton Bicknell, Jeremy Roschelle, Benjamin Motz 0002, Danielle S. McNamara, Richard G. Baraniuk, Debshila Basu Mallick, René F. Kizilcec, Ryan Baker 0001, Stephen Fancsali, April Murphy |
L@S | 3 |
| 2023 | Exploring The Potential of Chatbots to Provide Mental Well-being Support for Computer Science StudentsabstractComputer Science students are affected by a number of stressors, such as competition, which make it difficult for them to manage their mental well-being and mood. Students are often reluctant to use existing resources for support because they are difficult to access or perceived as ineffective. Conversational agents have shown potential to provide accessible and effective support to improve well-being. In this work, we explore the problem space to identify contexts in which chatbots could be beneficial for students and investigate how different types of chatbot could supplement existing resources provided by universities. Kunzhi Yu, Jiakai Shi, Joseph Jay Williams |
SIGCSE (2) | 5 |
| 2023 | A Case Study in Opportunities for Adaptive Experiments to Enable Rapid Continuous ImprovementabstractDrawing inspiration from machine learning and experimentation in product development at leading technology companies, we explore how adaptive experimentation might help in continuous course improvement. In adaptive experiments, as different arms/conditions are deployed to students, data is analyzed and used to change the experience for future students. We discuss an example side-by-side comparison of traditional and adaptive experimentation of self-explanation prompts in online homework problems in a CS1 course. This provides the first step in exploring the future of how this approach can help bridge research and practice in continuous course improvement. Ilya Musabirov, Angela M. Zavaleta Bernuy, Michael Liut, Joseph Jay Williams |
SIGCSE (2) | 4 |
| 2023 | Investigating Subject Lines Length on Students' Email Open RatesabstractInstructors often prefer to use email for course communication. The use of emails has been widely discussed in the fields of marketing and behavioural design, but the prevalence of email in education makes it important for instructors to collect metrics on emails to see how students engage with them. One component of emails are the subject lines, which constitute as one of the first things a receiver sees before deciding to open an email. This poster discusses a case study at deploying an email intervention in an online CS1 course. We investigate how the length of subject lines impact the rate at which students open emails of a particular type that prompts them to start their homework early. We aim to share key results to inform instructors how to design their emails to better reach students. Further, we highlight the potential benefits for instructors when collecting and analyzing email engagement data. Elexandra Tran, Angela M. Zavaleta Bernuy, Bogdan Simion, Michael Liut, Andrew Petersen 0001, Joseph Jay Williams |
SIGCSE (2) | 6 |
| 2023 | Designing, Deploying, and Analyzing Adaptive Educational Field ExperimentsabstractDigital experiments can be used in CSedu to test hypotheses about interventions and conditions' efficacy (or inefficacy). This workshop will discuss and deconstruct the design process and analysis for various experiments conducted in CS1. E.g., experiments testing which explanations students find helpful, which emails get them to start homework early, or which webpages effectively encourage and motivate students. This workshop teaches participants how to conduct, interpret, and analyze adaptive field experiments. These adaptive experiments employ machine learning algorithms to analyze experiments during deployment and dynamically shift the allocation of arms/conditions to give future students better conditions more rapidly. Adaptive field experiments can accelerate scientific discovery by enabling more complex experimental designs and increasing statistical power by phasing conditions in and out more efficiently. The workshop is supported by a 5-year NSF grant to build software tools and a digital community, gathering instructors, domain scientists and methodologists to teach them how to run adaptive experiments. The methodological focus includes understanding: (1) which algorithms are best for adaptive experiments that meet domain scientists' needs in specific experimental designs and data sets; (2) which hypothesis tests and Bayesian analyses to choose. Software companies use these innovative methodologies extensively to continuously improve product design. This workshop demonstrates how the same methods can be used in CSedu to improve research rigor and accelerate educational research implementation, ultimately improving student outcomes. Joseph Jay Williams, Nathan Laundry, Ilya Musabirov, Angela M. Zavaleta Bernuy, Michael Liut |
SIGCSE (2) | 1 |
| 2023 | Designing Voice Reflection for StudentsabstractResearch has revealed the positive effects of reflection on helping students manage their psychological well-being. Hence, we are motivated to investigate the design space of how we can communicate the values of doing reflections. We limit our work to voice reflection because of its inclusivity and effectiveness. We designed a pilot survey on Qualtrics to collect qualitative responses and integrated voice recording features from Phonic.ai. Participants were presented with 4 sample voice recordings related to college students' daily lives and asked to complete a simple voice reflection activity based on the samples they listened to. Then, they were asked to provide feedback on these examples. We deployed the survey on Amazon Mechanical Turk (MTurk) and collected 221 effective responses. By conducting thematic analysis, we found several insightful themes: emotional speech, diverse content, and clear structure are important elements to include, while examples should avoid being overly scripted. The findings suggest ways to design effective examples to engage students in voice reflections and open up the possibilities for further investigations into the design features of voice reflection platforms. Xuening Wu, Eunchae Seong, Ananya Bhattacharjee, Dana Kulzhabayeva, Pan Chen 0005, Joseph Jay Williams |
SIGCSE (2) | 6 |
| 2023 | Design Implications for One-Way Text Messaging Services that Support Psychological WellbeingabstractOne-way text messaging services have the potential to support psychological wellbeing at scale without conversational partners. However, there is limited understanding of what challenges are faced in mapping interactions typically done face-to-face or via online interactive resources into a text messaging medium. To explore this design space, we developed seven text messages inspired by cognitive behavioral therapy. We then conducted an open-ended survey with 788 undergraduate students and follow-up interviews with students and clinical psychologists to understand how people perceived these messages and the factors they anticipated would drive their engagement. We leveraged those insights to revise our messages, after which we deployed our messages via a technology probe to 11 students for two weeks. Through our mixed-methods approach, we highlight challenges and opportunities for future text messaging services, such as the importance of concrete suggestions and flexible pre-scheduled message timing. Ananya Bhattacharjee, Jiyau Pang, Angelina Liu, Alexander Mariakakis, Joseph Jay Williams |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2022 | Meeting Users Where They Are: User-centered Design of an Automated Text Messaging Tool to Support the Mental Health of Young AdultsabstractYoung adults have high rates of mental health conditions, but most do not want or cannot access formal treatment. We therefore recruited young adults with depression or anxiety symptoms to co-design a digital tool for self-managing their mental health concerns. Through study activities-consisting of an online discussion group and a series of design workshops-participants highlighted the importance of easy-to-use digital tools that allow them to exercise independence in their self-management. They described ways that an automated messaging tool might benefit them by: facilitating experimentation with diverse concepts and experiences; allowing variable depth of engagement based on preferences, availability, and mood; and collecting feedback to personalize the tool. While participants wanted to feel supported by an automated tool, they cautioned against incorporating an overtly human-like motivational tone. We discuss ways to apply these findings to improve the design and dissemination of digital mental health tools for young adults. Rachel Kornfield, Jonah Meyerhoff, Hannah Studd, Ananya Bhattacharjee, Joseph Jay Williams, Madhu C. Reddy, David C. Mohr |
CHI | 5 |
| 2022 | How can Email Interventions Increase Students' Completion of Online Homework? A Case Study Using A/B ComparisonsabstractEmail communication between instructors and students is ubiquitous, and it could be valuable to explore ways of testing out how to make email messages more impactful. This paper explores the design space of using emails to get students to plan and reflect on starting weekly homework earlier. We deployed a series of email reminders using randomized A/B comparisons to test alternative factors in the design of these emails, providing examples of an experimental paradigm and metrics for a broader range of interventions. We also surveyed and interviewed instructors and students to compare their predictions about the effectiveness of the reminders with their actual impact. We present our results on which seemingly obvious predictions about effective emails are not borne out, despite there being evidence for further exploring these interventions, as they can sometimes motivate students to attempt their homework more often. We also present qualitative evidence about student opinions and behaviours after receiving the emails, to guide further interventions. These findings provide insight into how to use randomized A/B comparisons in everyday channels such as emails, to provide empirical evidence to test our beliefs about the effectiveness of alternative design choices. Angela M. Zavaleta Bernuy, Ziwen Han, Hammad Shaikh, Qi Yin Zheng, Lisa-Angelique Lim, Anna N. Rafferty, Andrew Petersen 0001, Joseph Jay Williams |
LAK | 8 |
| 2022 | Third Annual Workshop on A/B Testing and Platform-Enabled Learning ResearchabstractLearning engineering adds tools and processes to learning platforms to support improvement research. One kind of tool is A/B testing, which is common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE). A number of A/B testing systems focused on educational applications have arisen recently, including UpGrade and E-TRIALS. A/B testing can be part of the puzzle of how to improve educational platforms, and yet challenging issues in education go beyond the generic paradigm. For example, the importance of teachers and instructors to learning means that students are not only connecting with software as individuals, but also as part of a shared classroom experience. Further, learning in topics like mathematics can be highly dependent on prior learning, and thus A or B may not be better overall, but only in interaction with prior knowledge. In response, a set of learning platforms is opening their systems to improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore how A/B testing in educational contexts is different, how learning platforms are opening up new possibilities, and how these empirical approaches can be used to drive powerful gains in student learning. It will also discuss forthcoming opportunities for funding to conduct platform-enabled learning research. Steven Ritter 0001, Neil T. Heffernan, Joseph Jay Williams, Derek Lomas, Benjamin Motz 0002, Debshila Basu Mallick, Klinton Bicknell, Danielle S. McNamara, René F. Kizilcec, Jeremy Roschelle, Richard G. Baraniuk, Ryan Baker 0001 |
L@S | 3 |
| 2022 | Investigating the Impact of Voice Response Options in SurveysabstractWith the widespread usage of mobile devices, users can now choose to provide input through voice or text. As researchers frequently ask students open-ended questions, we want to explore a natural mode to obtain better feedback in surveys. This study details a preliminary study demonstrating the importance of allowing students to choose between voice or text input to respond to surveys. A survey with several open-ended questions was deployed in a CS1 course. Correlations between the gender of the respondent and their method of responding were evaluated. We found that voice responses tended to be longer and preferred more by females relative to male students. Pan Chen 0005, Naaz Sibia, Angela M. Zavaleta Bernuy, Michael Liut, Joseph Jay Williams |
SIGCSE (2) | 5 |
| 2022 | "I Kind of Bounce off It": Translating Mental Health Principles into Real Life Through Story-Based Text MessagesabstractAdopting new psychological strategies to improve mental wellness can be challenging since people are often unable to anticipate how new habits are applicable to their circumstances. Narrative-based interventions have the potential to alleviate this burden by illustrating psychological principles in an applied context. In this work, we explore how stories can be delivered via the ubiquitous and scalable medium of text messaging. Through formative work consisting of interviews and focus group discussions with 15 participants, we identified desirable elements of stories about mental health, including authenticity and relatability. We then deployed story-based text messages to 42 participants to explore challenges regarding both the stories' content (e.g., specific versus generalized) and format (e.g., story length). We observed that our stories helped participants reflect on and identify flaws in their thinking patterns. Our findings highlight design implications and opportunities for mental wellness interventions that utilize stories in text messaging services. Ananya Bhattacharjee, Joseph Jay Williams, Karrie Chou, Justice Tomlinson, Jonah Meyerhoff, Alexander Mariakakis, Rachel Kornfield |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Involving Crowdworkers with Lived Experience in Content-Development for Push-Based Digital Mental Health Tools: Lessons Learned from Crowdsourcing Mental Health MessagesabstractDigital tools can support individuals managing mental health concerns, but delivering sufficiently engaging content is challenging. This paper seeks to clarify how individuals with mental health concerns can contribute content to improve push-based mental health messaging tools. We recruited crowdworkers with mental health symptoms to evaluate and revise expert-composed content for an automated messaging tool, and to generate new topics and messages. A second wave of crowdworkers evaluated expert and crowdsourced content. Crowdworkers generated topics for messages that had not been prioritized by experts, including self-care, positive thinking, inspiration, relaxation, and reassurance. Peer evaluators rated messages written by experts and peers similarly. Our findings also suggest the importance of personalization, particularly when content adaptation occurs over time as users interact with example messages. These findings demonstrate the potential of crowdsourcing for generating diverse and engaging content for push-based tools, and suggest the need to support users in meaningful content customization. Rachel Kornfield, David C. Mohr, Rachel Ranney, Emily G. Lattie, Jonah Meyerhoff, Joseph Jay Williams, Madhu C. Reddy |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2021 | Using Adaptive Experiments to Rapidly Help Students
Angela M. Zavaleta Bernuy, Qi Yin Zheng, Hammad Shaikh, Jacob Nogas, Anna N. Rafferty, Andrew Petersen 0001, Joseph Jay Williams |
AIED (2) | 7 |
| 2021 | Exploring Additional Personalized Support While Attempting Exercise Problems in Online Learning PlatformsabstractIn online asynchronous learning environments, students are assigned exercises, but it is not clear how to incorporate the kinds of actions an in-person tutor might take such as explaining, providing more practice, prompting for reflection, and motivating. We explore approaches to adding "Drop-Downs'' that appear after a student submits an answer and that contain additional information to support learning. We conducted randomized A/B experiments exploring the impact of these Drop-Downs on student learning in the online portion of a flipped CS1 course. The deployed Drop-Downs in this course provided explanations, reflective prompts, additional problems, and motivational messages. The results suggest that students benefit from various Drop-Downs in different contexts, indicating the possibility of personalizing content based on the student's state. We discuss the resulting design implications of Drop-Downs in online learning systems. Yuya Asano, Madhurima Dutta, Trisha Thakur, Jaemarie Solyst, Stephanie Cristea, Helena Jovic, Andrew Petersen 0001, Joseph Jay Williams |
L@S | 8 |
| 2021 | Exploring Design Choices in Data-driven Hints for Python Programming HomeworkabstractStudents often struggle during programming homework and may need help getting started or localizing errors. One promising and scalable solution is to provide automated programming hints, generated from prior student data, which suggest how a student can edit their code to get closer to a solution, but little work has explored how to design these hints for large-scale, real-world classroom settings, or evaluated such designs. In this paper, we present CodeChecker, a system which generates hints automatically using student data, and incorporates them into an existing CS1 online homework environment, used by over 1000 students per semester. We present insights from survey and interview data, about student and instructor perceptions of the system. Our results highlight affordances and limitations of automated hints, and suggest how specific design choices may have impacted their effectiveness. Thomas W. Price, Samiha Marwan, Joseph Jay Williams |
L@S | 3 |
| 2021 | The MOOClet Framework: Unifying Experimentation, Dynamic Improvement, and Personalization in Online CoursesabstractHow can educational platforms be instrumented to accelerate the use of research to improve students' experiences? We show how modular components of any educational interface - e.g. explanations, homework problems, even emails - can be implemented using the novel MOOClet software architecture. Researchers and instructors can use these augmented MOOClet components for: (1) Iterative Cycles of Randomized Experiments that test alternative versions of course content; (2) Data-Driven Improvement using adaptive experiments that rapidly use data to give better versions of content to future students, on the order of days rather than months. A MOOClet supports both manual and automated improvement using reinforcement learning; (3) Personalization by delivering alternative versions as a function of data about a student's characteristics or subgroup, using both expert-authored rules and data mining algorithms. We provide an open-source web service for implementing MOOClets (www.mooclet.org) that has been used with thousands of students. The MOOClet framework provides an ecosystem that transforms online course components into collaborative micro-laboratories, where instructors, experimental researchers, and data mining/machine learning researchers can engage in perpetual cycles of experimentation, improvement, and personalization. Mohi Reza, Juho Kim 0001, Ananya Bhattacharjee, Anna N. Rafferty, Joseph Jay Williams |
L@S | 5 |
| 2021 | Second Workshop on Educational A/B Testing at ScaleabstractThe emerging discipline of Learning Engineering is focused on putting into place tools and processes that use the science of learning as a basis for improving educational outcomes. An important part of Learning Engineering focuses on improving the effectiveness of educational software. In many software domains, A/B testing has become a prominent technique to achieve the software's goals. Many large companies (Amazon, Google, Facebook, etc.) run thousands of AB tests and present at the Annual Conference on Digital Experimentation (CODE), but that venue is too broad to address AB testing issues specific to EdTech platforms. We see a need to address issues with running large-scale A/B tests within the educational context, where the use of A/B testing lags other industries. This workshop will explore ways in which A/B testing in educational contexts differs from other domains and proposals to overcome current challenges so that this approach can become a more useful tool in the learning engineer's toolbox. Steven Ritter 0001, Neil T. Heffernan, Joseph Jay Williams, Derek Lomas, Klinton Bicknell |
L@S | 3 |
| 2021 | Investigating the Impact of Online Homework Reminders Using Randomized A/B ComparisonsabstractProcrastination by students may lead to adverse outcomes such as a focus on completion rather than learning or even a failure to complete learning tasks. One common method for motivating students and reducing procrastination is to send reminders with hints and study strategies, but it's not clear if these messages are effective or when is the best time to send them. Randomized A/B comparisons could be used to try different reminders or alternative ideas about how best to get students to start work earlier and, crucially, to measure the impact of these interventions on behaviour. This paper describes an A/B comparison of reminder emails set in a large CS1 course at a research-focused North American university. We found evidence that the email interventions caused a higher proportion of students to attempt the online homework but did not see evidence that these particular emails got students to start early, irrespective of changes to the timing of the reminder. More broadly, these findings illustrate how to use A/B comparisons in educational settings to test ideas about how to help students, and demonstrate the value of using randomized A/B comparisons, even when evaluating actions that seem obviously beneficial, such as reminder emails. Angela M. Zavaleta Bernuy, Qi Yin Zheng, Hammad Shaikh, Andrew Petersen 0001, Joseph Jay Williams |
SIGCSE | 5 |
| 2021 | Procrastination and Gaming in an Online Homework System of an Inverted CS1abstractEngaged preparation and study in combination with lectures are important for all courses but are particularly critical for online, hybrid, and inverted classrooms. Many instructors use online systems to deliver new course content and exercises, but students often delay assignments or game these systems (e.g., guessing on multiple-choice questions), often to the detriment of their learning. In an inverted CS1 course, many students self-reported high rates of gaming-the-system behavior, so we examine survey data to identify factors that contribute to engagement in these maladaptive behaviours. We supplement that analysis with interview data to gain a deeper understanding of the situation. We also implemented and evaluated a previously reported online intervention aimed at reducing gaming behavior. Unlike prior work, our intervention did not have a significant effect on guessing behavior. We discuss why the factors we identified might explain this result, as well as suggest future work to improve our understanding of gaming behaviours and inform the design of systems that encourage effective learning. Jaemarie Solyst, Trisha Thakur, Madhurima Dutta, Yuya Asano, Andrew Petersen 0001, Joseph Jay Williams |
SIGCSE | 6 |
| 2021 | Adaptive learning algorithms to optimize mobile applications for behavioral health: guidelines for design decisionsabstractOBJECTIVE: Providing behavioral health interventions via smartphones allows these interventions to be adapted to the changing behavior, preferences, and needs of individuals. This can be achieved through reinforcement learning (RL), a sub-area of machine learning. However, many challenges could affect the effectiveness of these algorithms in the real world. We provide guidelines for decision-making. MATERIALS AND METHODS: Using thematic analysis, we describe challenges, considerations, and solutions for algorithm design decisions in a collaboration between health services researchers, clinicians, and data scientists. We use the design process of an RL algorithm for a mobile health study "DIAMANTE" for increasing physical activity in underserved patients with diabetes and depression. Over the 1.5-year project, we kept track of the research process using collaborative cloud Google Documents, Whatsapp messenger, and video teleconferencing. We discussed, categorized, and coded critical challenges. We grouped challenges to create thematic topic process domains. RESULTS: Nine challenges emerged, which we divided into 3 major themes: 1. Choosing the model for decision-making, including appropriate contextual and reward variables; 2. Data handling/collection, such as how to deal with missing or incorrect data in real-time; 3. Weighing the algorithm performance vs effectiveness/implementation in real-world settings. CONCLUSION: The creation of effective behavioral health interventions does not depend only on final algorithm performance. Many decisions in the real world are necessary to formulate the design of problem parameters to which an algorithm is applied. Researchers must document and evaulate these considerations and decisions before and during the intervention period, to increase transparency, accountability, and reproducibility. TRIAL REGISTRATION: clinicaltrials.gov, NCT03490253. Caroline A. Figueroa, Adrián Aguilera, Bibhas Chakraborty, Arghavan Modiri, Jai Aggarwal, Nina Deliu, Urmimala Sarkar, Joseph Jay Williams, Courtney R. Lyles |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | Bandit algorithms to personalize educational chatbotsabstractTo emulate the interactivity of in-person math instruction, we developed MathBot, a rule-based chatbot that explains math concepts, provides practice questions, and offers tailored feedback. We evaluated MathBot through three Amazon Mechanical Turk studies in which participants learned about arithmetic sequences. In the first study, we found that more than 40% of our participants indicated a preference for learning with MathBot over videos and written tutorials from Khan Academy. The second study measured learning gains, and found that MathBot produced comparable gains to Khan Academy videos and tutorials. We solicited feedback from users in those two studies to emulate a real-world development cycle, with some users finding the lesson too slow and others finding it too fast. We addressed these concerns in the third and main study by integrating a contextual bandit algorithm into MathBot to personalize the pace of the conversation, allowing the bandit to either insert extra practice problems or skip explanations. We randomized participants between two conditions in which actions were chosen uniformly at random (i.e., a randomized A/B experiment) or by the contextual bandit. We found that the bandit learned a similarly effective pedagogical policy to that learned by the randomized A/B experiment while incurring a lower cost of experimentation. Our findings suggest that personalized conversational agents are promising tools to complement existing online resources for math education, and that data-driven approaches such as contextual bandits are valuable tools for learning effective personalization. William Cai, Joshua Grossman, Zhiyuan Jerry Lin, Hao Sheng 0003, Johnny Tian-Zheng Wei, Joseph Jay Williams, Sharad Goel |
Mach. Learn. | 6 |
| 2021 | Can Crowds Customize Instructional Materials with Minimal Expert Guidance?: Exploring Teacher-guided Crowdsourcing for Improving Hints in an AI-based TutorabstractAI-based educational technologies may be most welcome in classrooms when they align with teachers' goals, preferences, and instructional practices. Teachers, however, have scarce time to make such customizations themselves. How might the crowd be leveraged to help time-strapped teachers? Crowdsourcing pipelines have traditionally focused on content generation. It is an open question how a pipeline might be designed so the crowd can succeed in a revision/customization task. In this paper, we explore an initial version of a teacher-guided crowdsourcing pipeline designed to improve the adaptive math hints of an AI-based tutoring system so they fit teachers' preferences, while requiring minimal expert guidance. In two experiments involving 144 math teachers and 481 crowdworkers, we found that such an expert-guided revision pipeline could save experts' time and produce better crowd-revised hints (in terms of teacher satisfaction) than two comparison conditions. The revised hints however, did not improve on the existing hints in the AI tutor, which were carefully-written but still have room for improvement and customization. Further analysis revealed that the main challenge for crowdworkers may lie in understanding teachers' brief written comments and implementing them in the form of effective edits, without introducing new problems. We also found that teachers preferred their own revisions over other sources of hints, and exhibited varying preferences for hints. Overall, the results confirm that there is a clear need for customizing hints to individual teachers' preferences. They also highlight the need for more elaborate scaffolds so the crowd can have specific knowledge of the requirements that teachers have for hints. The study represents a first exploration in the literature of how to support crowds with minimal expert guidance in revising and customizing instructional materials. Kexin Bella Yang, Tomohiro Nagashima, Junhui Yao, Joseph Jay Williams, Kenneth Holstein, Vincent Aleven |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2020 | An Evaluation of Data-Driven Programming Hints in a Classroom Setting
Thomas W. Price, Samiha Marwan, Michael Winters, Joseph Jay Williams |
AIED (2) | 4 |
| 2020 | Challenges and opportunities of using reinforcement learning to optimize behavioral health interventions delivered via smartphones
Caroline A. Figueroa, Adrián Aguilera, Bibhas Chakraborty, Arghavan Modiri, Jai Aggarwal, Nina Deliu, Urmimala Sarkar, Joseph Jay Williams, Courtney R. Lyles |
AMIA | 8 |
| 2020 | SleepBandits: Guided Flexible Self-Experiments for SleepabstractSelf-experiments allow people to explore what behavioral changes lead to improved health and wellness. However, it is challenging to run such experiments in a scientifically valid way that is also flexible and able to accommodate the realities of daily life. We present a set of design principles for guided self-experiments that aim to lower this barrier to self-experimentation. We demonstrate the value of the principles by implementing them in SleepBandits, an integrated system that includes a smartphone application for sleep experiments. SleepBandits guides users through the steps of a single-case experiment, automatically collecting data from the built-in sensors and user input and calculating and presenting results in real-time. We released SleepBandits to the Google Play Store and people voluntarily downloaded and used it. Based on the data from 365 active users from this in-the-wild study, we discuss opportunities and challenges with the design principles and the SleepBandits system. Nediyana Daskalova, Jina Yoon, Cintia Araújo, Guillermo Beltrán, Nicole Nugent, John McGeary, Joseph Jay Williams, Jeff Huang 0002 |
CHI | 8 |
| 2020 | Engaging Students with Instructor Solutions in Online Programming HomeworkabstractStudents working on programming homework do not receive the same level of support as in the classroom, relying primarily on automated feedback from test cases. One low-effort way to provide more support is by prompting students to compare their solution to an instructor's solution, but it is unclear the best way to design such prompts to support learning. We designed and deployed a randomized controlled trial during online programming homework, where we provided students with an instructor's solution, and randomized whether they were prompted to compare their solution to the instructor's, to fill in the blanks for a written explanation of the instructor's solution, to do both, or neither. Our results suggest that these prompts can effectively engage students in reflecting on instructor solutions, although the results point to design trade-offs between the amount of effort that different prompts require from students and instructors, and their relative impact on learning. Thomas W. Price, Joseph Jay Williams, Jaemarie Solyst, Samiha Marwan |
CHI | 2 |
| 2020 | Getting too personal(ized): The importance of feature choice in online adaptive algorithms
Zhaobin Li, Luna Yee, Nathaniel Sauerberg, Irene Sakson, Joseph Jay Williams, Anna N. Rafferty |
EDM | 5 |
| 2020 | Characterizing and influencing students' tendency to write self-explanations in online homeworkabstractIn the context of online programming homework for a university course, we explore the extent to which learners engage with optional prompts to self -explain answers they choose for problems. Such prompts are known to benefit learning in laboratory and classroom settings [4], but there are less data about the extent to which students engage with them when they are optional additions to online homework. We report data from a deployment of self-explanation prompts in online programming homework, providing insight into how the frequency of writing explanations is correlated with different variables, such as how early students start homework, whether they got a problem correct, and how proficient they are in the language of instruction. We also report suggestive results from a randomized experiment comparing several methods for increasing the rate at which people write explanations, such as including more than one kind of prompt. These findings provide insight into promising dimensions to explore in understanding how real students may engage with prompts to explain answers. Yuya Asano, Jaemarie Solyst, Joseph Jay Williams |
LAK | 3 |
| 2020 | Workshop Proposal: Educational A/B Testing at ScaleabstractNo abstract available. Steven Ritter 0001, Neil T. Heffernan, Joseph Jay Williams, Burr Settles, Phillip Grimaldi, Derek Lomas |
L@S | 3 |
| 2020 | Using Information Visualization to Promote Students' Reflection on "Gaming the System" in Online Learningabstract"Gaming the system" is the phenomenon where students attempt to perform well by systematically exploiting properties of the learning system, rather than learning the material. Frequent gaming tends to cause bad learning outcomes. Though existing studies tackle the problem by redesigning the system workflow to change students' behaviors automatically, gaming students discover new ways to game. We instead propose a novel way, reflective nudge, to reflectively influence students' attitudes by conveying reasons not to game via information visualizations. Particularly, we identify three common gaming contexts and involve students and instructors in co-designing three context-specific persuasive visualizations. We deploy our information visualizations in a real online learning platform. Through embedded surveys and in-person interviews, we find some evidence that the designs can promote students' reflection on gaming, and suggestive data that two of them can reduce gaming compared with control groups. Furthermore, we present insights into reflective nudge designs and practical issues concerning deployment. Meng Xia 0002, Yuya Asano, Joseph Jay Williams, Huamin Qu, Xiaojuan Ma |
L@S | 3 |
| 2019 | Balancing Student Success and Inferring Personalized Effects in Dynamic Experiments
Hammad Shaikh, Arghavan Modiri, Joseph Jay Williams, Anna N. Rafferty |
EDM | 3 |
| 2019 | An Evaluation of the Impact of Automated Programming Hints on Performance and LearningabstractA growing body of work has explored how to automatically generate hints for novice programmers, and many programming environments now employ these hints. However, few studies have investigated the efficacy of automated programming hints for improving performance and learning, how and when novices find these hints beneficial, and the tradeoffs that exist between different types of hints. In this work, we explored the efficacy of next-step code hints with 2 complementary features: textual explanations and self-explanation prompts. We conducted two studies in which novices completed two programming tasks in a block-based programming environment with automated hints. In Study 1, 10 undergraduate students completed 2 programming tasks with a variety of hint types, and we interviewed them to understand their perceptions of the affordances of each hint type. For Study 2, we recruited a convenience sample of participants without programming experience from Amazon Mechanical Turk. We conducted a randomized experiment comparing the effects of hints' types on learners' performance and performance on a subsequent task without hints. We found that code hints with textual explanations significantly improved immediate programming performance. However, these hints only improved performance in a subsequent post-test task with similar objectives, when they were combined with self-explanation prompts. These results provide design insights into how automatically generated code hints can be improved with textual explanations and prompts to self-explain, and provide evidence about when and how these hints can improve programming performance and learning. Samiha Marwan, Joseph Jay Williams, Thomas W. Price |
ICER | 2 |
| 2019 | Learning Multi-Objective Rewards and User Utility Function in Contextual Bandits for Personalized RankingabstractThis paper tackles the problem of providing users with ranked lists of relevant search results, by incorporating contextual features of the users and search results, and learning how a user values multiple objectives. For example, to recommend a ranked list of hotels, an algorithm must learn which hotels are the right price for users, as well as how users vary in their weighting of price against the location. In our paper, we formulate the context-aware, multi-objective, ranking problem as a Multi-Objective Contextual Ranked Bandit (MOCR-B). To solve the MOCR-B problem, we present a novel algorithm, named Multi-Objective Utility-Upper Confidence Bound (MOU-UCB). The goal of MOU-UCB is to learn how to generate a ranked list of resources that maximizes the rewards in multiple objectives to give relevant search results. Our algorithm learns to predict rewards in multiple objectives based on contextual information (combining the Upper Confidence Bound algorithm for multi-armed contextual bandits with neural network embeddings), as well as learns how a user weights the multiple objectives. Our empirical results reveal that the ranked lists generated by MOU-UCB lead to better click-through rates, compared to approaches that do not learn the utility function over multiple reward objectives. Nirandika Wanigasekara, Yuxuan Liang 0002, Siong-Thye Goh, Ye Liu 0002, Joseph Jay Williams, David S. Rosenblum |
IJCAI | 5 |
| 2019 | The Impact of Adding Textual Explanations to Next-step Hints in a Novice Programming EnvironmentabstractAutomated hints, a powerful feature of many programming environments, have been shown to improve students' performance and learning. New methods for generating these hints use historical data, allowing them to scale easily to new classrooms and contexts. These scalable methods often generate next-step, code hints that suggest a single edit for the student to make to their code. However, while these code hints tell the student what to do, they do not explain why, which can make these hints hard to interpret and decrease students' trust in their helpfulness. In this work, we augmented code hints by adding adaptive, textual explanations in a block-based, novice programming environment. We evaluated their impact in two controlled studies with novice learners to investigate how our results generalize to different populations. We measured the impact of textual explanations on novices' programming performance. We also used quantitative analysis of log data, self-explanation prompts, and frequent feedback surveys to evaluate novices' understanding and perception of the hints throughout the learning process. Our results showed that novices perceived hints with explanations as significantly more relevant and interpretable than those without explanations, and were also better able to connect these hints to their code and the assignment. However, we found little difference in novices' performance. Our results suggest that explanations have the potential to make code hints more useful, but it is unclear whether this translates into better overall performance and learning. Samiha Marwan, Nicholas Lytle, Joseph Jay Williams, Thomas W. Price |
ITiCSE | 3 |
| 2018 | Bandit Assignment for Educational Experiments: Benefits to Students Versus Statistical Power
Anna N. Rafferty, Huiji Ying, Joseph Jay Williams |
AIED (2) | 3 |
| 2018 | Combining Difficulty Ranking with Multi-Armed Bandits to Sequence Educational Content
Avi Segal, Yossi Ben David, Joseph Jay Williams, Kobi Gal, Yaar Shalom |
AIED (2) | 3 |
| 2018 | Harvesting Caregiving Knowledge: Design Considerations for Integrating Volunteer Input in Dementia CareabstractImproving volunteer performance leads to better caregiving in dementia care settings. However, caregiving knowledge systems have been focused on eliciting and sharing expert, primary caregiver knowledge, rather than volunteer-provided knowledge. Through the use of an experience prototype, we explored the content of volunteer caregiver knowledge and identified ways in which such non-expert knowledge can be useful to dementia care. By using lay language, sharing information specific to the client and collaboratively finding strategies for interaction, volunteers were able to boost the effectiveness of future volunteers. Therapists who reviewed the content affirmed the reliability of volunteer caregiver knowledge and placed value on its recency, variety and its ability to help bridge language and professional barriers. We discuss how future systems designed for eliciting and sharing volunteer caregiver knowledge can be used to promote better dementia care. Pin Sym Foong, Shengdong Zhao 0001, Felicia Fang-Yi Tan, Joseph Jay Williams |
CHI | 4 |
| 2018 | Understanding the Effect of In-Video Prompting on Learners and InstructorsabstractOnline instructional videos are ubiquitous, but it is difficult for instructors to gauge learners' experience and their level of comprehension or confusion regarding the lecture video. Moreover, learners watching the videos may become disengaged or fail to reflect and construct their own understanding. This paper explores instructor and learner perceptions of in-video prompting where learners answer reflective questions while watching videos. We conducted two studies with crowd workers to understand the effect of prompting in general, and the effect of different prompting strategies on both learners and instructors. Results show that some learners found prompts to be useful checkpoints for reflection, while others found them distracting. Instructors reported the collected responses to be generally more specific than what they have usually collected. Also, different prompting strategies had different effects on the learning experience and the usefulness of responses as feedback. Hyungyu Shin, Eun-Young Ko, Joseph Jay Williams, Juho Kim 0001 |
CHI | 3 |
| 2018 | Enhancing Online Problems Through Instructor-Centered Tools for Randomized ExperimentsabstractDigital educational resources could enable the use of randomized experiments to answer pedagogical questions that instructors care about, taking academic research out of the laboratory and into the classroom. We take an instructor-centered approach to designing tools for experimentation that lower the barriers for instructors to conduct experiments. We explore this approach through DynamicProblem, a proof-of-concept system for experimentation on components of digital problems, which provides interfaces for authoring of experiments on explanations, hints, feedback messages, and learning tips. To rapidly turn data from experiments into practical improvements, the system uses an interpretable machine learning algorithm to analyze students' ratings of which conditions are helpful, and present conditions to future students in proportion to the evidence they are higher rated. We evaluated the system by collaboratively deploying experiments in the courses of three mathematics instructors. They reported benefits in reflecting on their pedagogy, and having a new method for improving online problems for future students. Joseph Jay Williams, Anna N. Rafferty, Dustin Tingley, Andrew M. Ang, Walter S. Lasecki, Juho Kim 0001 |
CHI | 1 |
| 2017 | MOOClets: A Framework for Dynamic Experimentation and PersonalizationabstractRandomized experiments in online educational environments are ubiquitous as a scientific method for investigating learning and motivation, but too rarely improve educational resources and produce practical benefits for learners. We suggest that software and tools for experimentally comparing resources are designed primarily through the lens of experiments as a scientific methodology, and therefore miss a tremendous opportunity for online experiments to serve as engines for dynamic improvement and personalization. We present the MOOClet requirements specification to guide the implementation of software or tools for experiments to ensure that whenever alternative versions of a resource can be experimentally compared (by randomly assigning versions), the resource can also be dynamically improved (by changing which versions are presented), and personalized (by presenting different versions to different people). The MOOClet specification was used to implement DEXPER, a proof-of-concept web service backend that enables dynamic experimentation and personalization of resources embedded in front-end educational platforms. We describe three use cases of MOOClets for dynamic experimentation and personalization of motivational emails, explanations, and problems. Joseph Jay Williams, Anna N. Rafferty, Samuel G. Maldonado, Andrew M. Ang, Dustin Tingley, Juho Kim 0001 |
L@S | 1 |
| 2017 | Educational Question Routing in Online Student CommunitiesabstractStudents' performance in Massive Open Online Courses (MOOCs) is enhanced by high quality discussion forums or recently emerging educational Community Question Answering (CQA) systems. Nevertheless, only a small number of students answer questions asked by their peers. This results in instructor overload, and many unanswered questions. To increase students' participation, we present an approach for recommendation of new questions to students who are likely to provide answers. Existing approaches to such question routing proposed for non-educational CQA systems tend to rely on a few experts, what is not applicable in educational domain where it is important to involve all kinds of students. In tackling this novel educational question routing problem, our method (1) goes beyond previous question-answering data as it incorporates additional non-QA data from the course (to improve prediction accuracy and to involve more of the student community) and (2) applies constraints on users' workload (to prevent user overloading). We use an ensemble classifier for predicting students' willingness to answer a question, as well as students' expertise for answering. We conducted an online evaluation of the proposed method using an A/B experiment in our CQA system deployed in edX MOOC. The proposed method outperformed a baseline method (non-educational question routing enhanced with workload restriction) by improving recommendation accuracy, keeping more community members active, and increasing an average number of their contributions. Jakub Macina, Ivan Srba, Joseph Jay Williams, Mária Bieliková |
RecSys | 3 |
| 2016 | Revising Learner Misconceptions Without Feedback: Prompting for Reflection on AnomaliesabstractThe Internet has enabled learning at scale, from Massive Open Online Courses (MOOCs) to Wikipedia. But online learners may become passive, instead of actively constructing knowledge and revising their beliefs in light of new facts. Instructors cannot directly diagnose thousands of learners' misconceptions and provide remedial tutoring. This paper investigates how instructors can prompt learners to reflect on facts that are anomalies with respect to their existing misconceptions, and how to choose these anomalies and prompts to guide learners to revise incorrect beliefs without any feedback. We conducted two randomized experiments with online crowd workers learning statistics. Results show that prompts to explain why these anomalies are true drive revision towards correct beliefs. But prompts to simply articulate thoughts about anomalies have no effect on learning. Furthermore, we find that explaining multiple anomalies is more effective than explaining only one, but the anomalies should rule out multiple misconceptions simultaneously. Joseph Jay Williams, Tania Lombrozo, Anne Hsu, Bernd Huber, Juho Kim 0001 |
CHI | 1 |
| 2016 | Discovering 'Tough Love' Interventions Despite Dropout
Joseph Jay Williams, Anthony Botelho, Adam Sales, Neil T. Heffernan, Charles Lang |
EDM | 1 |
| 2016 | The assessment of learning infrastructure (ALI): the theory, practice, and scalability of automated assessmentabstractResearchers invested in K-12 education struggle not just to enhance pedagogy, curriculum, and student engagement, but also to harness the power of technology in ways that will optimize learning. Online learning platforms offer a powerful environment for educational research at scale. The present work details the creation of an automated system designed to provide researchers with insights regarding data logged from randomized controlled experiments conducted within the ASSISTments TestBed. The Assessment of Learning Infrastructure (ALI) builds upon existing technologies to foster a symbiotic relationship beneficial to students, researchers, the platform and its content, and the learning analytics community. ALI is a sophisticated automated reporting system that provides an overview of sample distributions and basic analyses for researchers to consider when assessing their data. ALI's benefits can also be felt at scale through analyses that crosscut multiple studies to drive iterative platform improvements while promoting personalized learning. Korinn S. Ostrow, Douglas Selent, Yan Wang 0005, Eric Van Inwegen, Neil T. Heffernan, Joseph Jay Williams |
LAK | 6 |
| 2016 | AXIS: Generating Explanations at Scale with Learnersourcing and Machine LearningabstractWhile explanations may help people learn by providing information about why an answer is correct, many problems on online platforms lack high-quality explanations. This paper presents AXIS (Adaptive eXplanation Improvement System), a system for obtaining explanations. AXIS asks learners to generate, revise, and evaluate explanations as they solve a problem, and then uses machine learning to dynamically determine which explanation to present to a future learner, based on previous learners' collective input. Results from a case study deployment and a randomized experiment demonstrate that AXIS elicits and identifies explanations that learners find helpful. Providing explanations from AXIS also objectively enhanced learning, when compared to the default practice where learners solved problems and received answers without explanations. The rated quality and learning benefit of AXIS explanations did not differ from explanations generated by an experienced instructor. Joseph Jay Williams, Juho Kim 0001, Anna N. Rafferty, Samuel G. Maldonado, Krzysztof Z. Gajos, Walter S. Lasecki, Neil T. Heffernan |
L@S | 1 |
| 2015 | Beyond Prediction: Towards Automatic Intervention in MOOC Student Stop-out
Jacob Whitehill, Joseph Jay Williams, Glenn Lopez, Cody A. Coleman, Justin Reich |
EDM | 2 |
| 2015 | A Playful Game Changer: Fostering Student Retention in Online Education with Social GamificationabstractMany MOOCs report high drop off rates for their students. Among the factors reportedly contributing to this picture are lack of motivation, feelings of isolation, and lack of interactivity in MOOCs. This paper investigates the potential of gamification with social game elements for increasing retention and learning success. Students in our experiment showed a significant increase of 25% in retention period (videos watched) and 23% higher average scores when the course interface was gamified. Social game elements amplify this effect significantly -- students in this condition showed an increase of 50% in retention period and 40% higher average test scores. Markus Krause, Marc Mogalle, Henning Pohl, Joseph Jay Williams |
L@S | 4 |
| 2015 | Supporting Instructors in Collaborating with Researchers using MOOCletsabstractMost education and workplace learning takes place in classroom contexts far removed from laboratories or field sites with special arrangements for scientific research. But digital online resources provide a novel opportunity for large-scale efforts to bridge the real-world and laboratory settings which support data collection and randomized A/B experiments comparing different versions of content or interactions [2]. However, there are substantial technological and practical barriers in aligning instructors and researchers to use learning technologies like blended lessons/exercises & MOOCs as both a service for students and a realistic context to conduct research. This paper explains how the concept of a "MOOClet" can facilitate research-practitioner collaborations. MOOClets [3] are defined as modular components of a digital resource that can be implemented in technology to: (1) allow modification to create multiple versions, (2) allow experimental comparison and personalization of different versions, (3) reliably specify what data are collected. We suggest a framework in which instructors specify what kinds of changes to lessons, exercises, and emails they would be willing to adopt, and what data they will collect and make available. Researchers can then: (1) specify or design experiments that compare the effects of different versions on quantifiable outcomes. (2) Explore algorithms for maximizing particular outcomes by choosing alternative versions of a MOOClet based on the input variables available. We present a prototype survey tool for instructors intended to facilitate practitioner-researcher matches and successful collaborations. Joseph Jay Williams, Juho Kim 0001, Brian Keegan |
L@S | 1 |
| 2015 | Using and Designing Platforms for In Vivo Educational ExperimentsabstractIn contrast to typical laboratory experiments, the everyday use of online educational resources by large populations and the prevalence of software infrastructure for A/B testing leads us to consider how platforms can embed in vivo experiments that do not merely support research, but ensure practical improvements to their educational components. Examples are presented of randomized experimental comparisons conducted by subsets of the authors in three widely used online educational platforms -- Khan Academy, edX, and ASSISTments. We suggest design principles for platform technology to support randomized experiments that lead to practical improvements -- enabling Iterative Improvement and Collaborative Work -- and explain the benefit of their implementation by WPI co-authors in the ASSISTments platform. Joseph Jay Williams, Korinn S. Ostrow, Xiaolu Xiong, Elena L. Glassman, Juho Kim 0001, Samuel G. Maldonado, Na Li 0002, Justin Reich, Neil T. Heffernan |
L@S | 1 |
| 2014 | Effects of Comparison and Explanation on Analogical Transfer
Brian J. Edwards, Joseph Jay Williams, Dedre Gentner, Tania Lombrozo |
CogSci | 2 |
| 2013 | Effects of Explanation and Comparison on Category Learning
Brian J. Edwards, Joseph Jay Williams, Tania Lombrozo |
CogSci | 2 |
| 2013 | Online Education: A Unique Opportunity for Cognitive Scientists to Integrate Research and Practice
Joseph Jay Williams, Alexander Renkl, Kenneth R. Koedinger, John C. Stamper |
CogSci | 1 |
| 2013 | Effects of Explaining Anomalies on the Generation and Evaluation of Hypotheses
Joseph Jay Williams, Caren M. Walker, Samuel G. Maldonado, Tania Lombrozo |
CogSci | 1 |
| 2013 | Evaluating computational models of explanation using human judgments
Michael Pacer, Joseph Jay Williams, Tania Lombrozo, Thomas L. Griffiths 0001 |
UAI | 2 |
| 2012 | Explaining Influences Children's Reliance on Evidence and Prior Knowledge in Causal Induction
Caren M. Walker, Joseph Jay Williams, Tania Lombrozo, Alison Gopnik |
CogSci | 2 |
| 2012 | Explaining increases belief revision in the face of (many) anomalies
Joseph Jay Williams, Caren M. Walker, Tania Lombrozo |
CogSci | 1 |
| 2011 | Explanation-based Mechanisms for Learning: An Interdisciplinary Approach
Michelene T. H. Chi, Gerald DeJong, Cristine H. Legare, Tania Lombrozo, Joseph Jay Williams |
CogSci | 5 |
| 2011 | Explaining drives the discovery of real and illusory patterns
Joseph Jay Williams, Tania Lombrozo, Bob Rehder |
CogSci | 1 |
| 2008 | Modeling human function learning with Gaussian processesabstractAccounts of how people learn functional relationships between continuous variables have tended to focus on two possibilities: that people are estimating explicit functions, or that they are simply performing associative learning supported by similarity. We provide a rational analysis of function learning, drawing on work on regression in machine learning and statistics. Using the equivalence of Bayesian linear regression and Gaussian processes, we show that learning explicit rules and using similarity can be seen as two views of one solution to this problem. We use this insight to define a Gaussian process model of human function learning that combines the strengths of both approaches. Thomas L. Griffiths 0001, Christopher G. Lucas, Joseph Jay Williams, Michael L. Kalish |
NIPS | 3 |