René F. Kizilcec

dblp:127/7170 · DBLP profile ↗
← Back
69ranked-venue papers
16as first author
45since 2021 · last 2026
0000-0001-6283-5546ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 57 · 10 first-author · 41 since 2021Artificial intelligence and machine learning · 38 · 8 first-author · 27 since 2021Systems, architecture and hardware · 38 · 8 first-author · 27 since 2021Human-computer interaction and ubiquitous computing · 27 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 Does the TalkMoves Codebook Generalize to One-on-One Tutoring and Multimodal Interaction?
Corina Luca Focsan, Marie Cynthia Abijuru Kamikazi, Tamisha Thompson, Jennifer St. John, Kirk Vanacore, Danielle R. Thomas, Kenneth R. Koedinger, René F. Kizilcec
AIED (5)8
2026 Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
Danielle R. Thomas, Conrad Borchers, Kirk Vanacore, Kenneth R. Koedinger, René F. Kizilcec
AIED (6)5
2026 AI Annotation Orchestration: Evaluating LLM Verifiers to Improve the Quality of LLM Annotations in Learning Analytics
abstract
Large Language Models (LLMs) are increasingly used to annotate learning interactions, yet concerns about reliability limit their utility. We test whether verification-oriented orchestration-prompting models to check their own labels (self-verification) or audit one another (cross-verification)-improves qualitative coding of tutoring discourse. Using transcripts from 30 one-to-one math sessions, we compare three production LLMs (GPT, Claude, Gemini) under three conditions: unverified annotation, self-verification, and cross-verification across all orchestration configurations. Outputs are benchmarked against a blinded, disagreement-focused human adjudication using Cohen's kappa. Overall, orchestration yields a 58 percent improvement in kappa. Self-verification nearly doubles agreement relative to unverified baselines, with the largest gains for challenging tutor moves. Cross-verification achieves a 37 percent improvement on average, with pair- and construct-dependent effects: some verifier-annotator pairs exceed self-verification, while others reduce alignment, reflecting differences in verifier strictness. We contribute: (1) a flexible orchestration framework instantiating control, self-, and cross-verification; (2) an empirical comparison across frontier LLMs on authentic tutoring data with blinded human "gold" labels; and (3) a concise notation, verifier(annotator) (e.g., Gemini(GPT) or Claude(Claude)), to standardize reporting and make directional effects explicit for replication. Results position verification as a principled design lever for reliable, scalable LLM-assisted annotation in Learning Analytics.
Bakhtawar Ahtisham, Kirk Vanacore, Jinsook Lee, Zhuqian Zhou, Doug Pietrzak, René F. Kizilcec
LAK6
2026 MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education
abstract
Medical education faces challenges in providing scalable, consistent clinical skills training. Simulation with standardized patients (SPs) develops communication and diagnostic skills, but remains resource-intensive and variable in feedback quality. Existing AI-based tools show promise yet often lack comprehensive assessment frameworks, evidence of clinical impact, and integration of self-regulated learning (SRL) principles. Through a multi-phase co-design process with medical education experts, we developed MedSimAI, an AI-powered simulation platform that enables deliberate practice through interactive patient encounters with immediate, structured feedback. Leveraging large language models, MedSimAI generates realistic clinical interactions and provides automated assessments aligned with validated evaluation frameworks. In a multi-institutional deployment (410 students; 1,024 encounters across three medical schools), 59.5% engaged in repeated practice. At one site, mean Objective Structured Clinical Examination (OSCE) history-taking scores rose from 82.8 to 88.8 (p < 0.001, d = 0.75), while a second site’s pilot showed no significant change. Automated scoring achieved 87% accuracy in identifying proficiency thresholds on the Master Interview Rating Scale (MIRS). Mixed-effects analyses revealed institution and case effects. Thematic analysis of 840 learner reflections highlighted challenges in missed items, organization, review-of-systems, and empathy. These findings position MedSimAI as a scalable formative platform for history-taking and communication, motivating staged curriculum integration and realism enhancements for advanced learners.
Yann Hicke, Jadon Geathers, Kellen Vu, Justin Sewell, Claire Cardie, Jaideep Talwalkar, Dennis L. Shung, Anyanate Gwendolyne Jack, Susannah Cornes, MacKenzi Preston, René F. Kizilcec
LAK11
2026 Who Decides in AI-Mediated Learning? The Agency Allocation Framework
abstract
As AI-mediated learning systems increasingly shape how learners plan, make decisions, and progress through education, learner agency is becoming both more consequential and harder to conceptualize at scale. Existing research often treats agency as a proxy for engagement and self-regulation, leaving unclear who actually holds decision-making authority in large-scale, automated learning environments. This paper reframes learner agency as the allocation of decision authority across learners, educators, institutions, and AI systems. We introduce the Agency Allocation Framework (AAF) for analyzing how decisions are distributed, how choices are architected, what evidence supports them, and over what time horizons their consequences unfold. Drawing on a focused review of Learning at Scale literature and an illustrative tutoring-system example, we identify four recurring challenges for studying learner agency at scale: (1) conceptual ambiguity, (2) reliance on behavioral proxies, (3) trade-offs between efficiency and learner control, and (4) the redistribution of agency through AI-mediated systems. Rather than advocating more or less automation, the AAF supports systematic analysis of when AI scaffolds learners' capacity to act and when it substitutes for it. By making decision authority explicit, the framework provides researchers and designers with analytic tools for studying, comparing, and evaluating agency-preserving learning systems in increasingly automated educational contexts.
Conrad Borchers, Olga Viberg, René F. Kizilcec
L@S3
2026 The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
abstract
Large language models (LLMs) have become popular writing tools among students and may expand access to high-quality feedback for students with less access to traditional writing support. At the same time, LLMs may standardize student voice or invite overreliance. This study examines how adoption of LLM-assisted writing varies across socioeconomic groups and how it relates to outcomes in a high-stakes context: U.S. college admissions. We analyze a de-identified longitudinal dataset of applications to a selective university from 2020 to 2024 (N = 81,663). Estimating LLM use using a distribution-based detector trained on synthetic and historical essays, we tracked how student writing changed as LLM use proliferated, how adoption differed by socioeconomic status (SES), and whether potential benefits translated equitably into admissions outcomes. Using fee-waiver status as a proxy for SES, we observe post-2023 convergence in surface-level linguistic features, with the largest changes among lower SES and rejected applicants. Estimated LLM use rose sharply in 2024 across all groups, with disproportionately larger increases among lower SES applicants, consistent with the hypothesis that LLMs substitute for scarce writing support. However, increased estimated LLM use was associated with larger declines in predicted admission probability for lower SES applicants than for higher SES applicants, even after controlling for academic credentials and stylometric features. These findings raise concerns about equity and the validity of essay-based evaluation in an era of AI-assisted writing and provide the first large-scale longitudinal evidence linking LLM adoption, linguistic change, and evaluative outcomes in college admissions.
Jinsook Lee, Conrad Borchers, A. J. Alvero, Thorsten Joachims, René F. Kizilcec
L@S5
2026 A Large-Scale Analysis of Student Behavior with Pedagogically Constrained LLM Tutors
Chang Liu 0122, Loc Hoang, René F. Kizilcec, Bo Wu 0002
L@S3
2026 Comparing Teacher and AI-Generated Feedback in the Writing Classroom: Experimental Results from Secondary School Classrooms
abstract
Writing proficiency is an essential skill for secondary school students and can be supported through high-quality feedback. However, providing individualized feedback on student writing is timeintensive, which often limits its availability in classroom practice. Recent advances in large language models (LLMs) have raised interest in automated feedback as a potential way to scale and supplement teacher practice, yet evidence of its effectiveness relative to teacher feedback in authentic classroom settings remains limited. In this study, we examined the effects of LLM-generated feedback on students' essay revisions, subsequent writing performance, and feedback perceptions in an authentic English-as-a-foreignlanguage context in secondary schools (N = 391). We compared teacher-written feedback against both delayed and immediate LLMgenerated feedback. We further conducted Bayesian analyses to characterize the magnitude and uncertainty of differences between feedback conditions. Together, these findings point to a trade-off between small performance differences associated with teacher feedback that were neither statistically or practically significant. At the same time, results showed more positive perceptions of usefulness and motivation associated with immediate feedback, which only LLMs can provide at scale. These findings suggest that the primary educational value of LLM-based feedback lies less in outperforming teachers and more in redistributing instructional effort. When used to consistently support revision and short-term learning, AI feedback may free teachers to focus on higher-order instructional planning, targeted scaffolding, and pedagogical interaction, highlighting the potential of hybrid feedback systems in writing instruction.
Jennifer Meyer, Marlene Steinbach, Ronja Schiller, Ute Mertens, Nils-Jonathan Schaller, Andrea Horbach, Johanna Fleckenstein, Olaf Köller, René F. Kizilcec, Thorben Jansen
L@S9
2026 Teachers' Perceived Benefits and Risks of AI Across Fifty-Five Countries: An Audit of LLM Alignment and Steerability
abstract
Teachers' trust in artificial intelligence (AI) in education depends on how they balance its perceived benefits and risks. Yet global discussions about scaling AI in education rely on fragmented evidence, as most studies of teachers' perceptions focus on single countries or small samples. This lack of representative cross-national evidence limits both theory building and policy development. At the same time, large language models (LLMs) are increasingly used in research, policy, and teachers' professional workflows, despite limited validation in education. To address these gaps, we conduct a large-scale audit of LLM alignment with teachers' perceptions of AI by combining representative international survey data with systematic model evaluation. Using OECD TALIS data from 55 countries and territories, we measure cross-national variation in teachers' perceived benefits and risks of AI. We then benchmark responses from eight state-of-the-art LLMs across four providers under both general and country-specific prompting, comparing higher- and lower-reasoning models. Results reveal substantial cross-national variation in teacher perceptions that is not reliably reflected in LLM outputs. Models compress country differences, overestimate both benefits and risks, and show limited gains from identity prompting or enhanced reasoning. This misalignment matters because LLM-generated guidance and professional discourse increasingly shape how teachers learn about and discuss AI, potentially influencing trust and future adoption decisions. Our findings caution against treating LLM outputs as substitutes for direct engagement with teachers when informing global AI-in-education initiatives. At the same time, some models (e.g., Gemini 3 Fast) partially capture cross-national ranking patterns, suggesting a complementary role in hypothesis generation and exploratory comparative analysis.
Yan Tao, Olga Viberg, Deepak Varuvel Dennison, Zhikun Wu, René F. Kizilcec
L@S5
2026 How Well Do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
abstract
Large language models (LLMs) are increasingly used in educational contexts, yet their ability to interpret authentic instructional discourse out-of-the-box remains unclear. We benchmark six state-of-the-art LLMs on classifying instructional moves in K-12 mathematics classroom transcripts annotated by expert educators (κ>0.90). We evaluated four prompting strategies, including zero-shot, one-shot, and few-shot prompts derived from the human coding manual. Zero-shot prompting achieved fair-to-moderate agreement (κ = 0.38–0.48, F1 = 0.45–0.53). Providing comprehensive examples improved performance for some models (e.g., κ = 0.48 to 0.58 for Claude 4.5 Opus; κ = 0.38 to 0.57 for Gemini 2.5 Pro), but gains were uneven and precision remained limited (best precision = 0.56, recall = 0.75). Errors were concentrated in constructs that require inference about instructor intent; for example, models confused Press for Reasoning with Press for Accuracy (42%–53% false-positive rates). Overall, our analysis found that LLMs demonstrate meaningful but limited capacity to identify aspects of instructional discourse, providing a baseline for educational discourse benchmarking and for designing more reliable annotation workflows. This work also points to a potential weakness in LLMs' ability to interpret key nuances of educational instruction.
Kirk Vanacore, René F. Kizilcec
L@S2
2026 From Tutor Moves to Tutoring States: Modeling the Timing and Sequencing of Pedagogical Strategies for Student Engagement
abstract
Understanding how tutoring unfolds requires capturing not only what tutors and students say, but also the multi-turn pedagogical strategies that structure their conversation. In this study, we analyze 77 online math tutoring sessions (about 40 hours of tutoring) using a combination of LLM-assisted annotation and statistical modeling. Building on utterance-level annotations with a taxonomy of tutor moves, we identify three recurrent latent states using a Hidden Markov Model: Inquiry Elicitation (dominated by prompting and probing student reasoning), Direct Instruction (centered on explanation and scaffolding), and Affective Support (characterized by praise and socioemotional support). Sequential pattern mining revealed that these states exhibit distinct instructional motifs—for example, Inquiry Elicitation involves repeated prompting, while Direct Instruction features sustained explanatory scaffolding. These dynamics were correlated with different levels of student engagement. Sessions starting with Inquiry Elicitation and ending with Affective Support elicited more student talk, while sustaining Direct Instruction was associated with a higher likelihood that students met tutoring dosage recommendations by returning for additional sessions. Together, these results suggest that the impact of tutoring may lie not just in which strategies tutors employ, but in how those strategies are ordered and coordinated over time—patterns that reveal the deeper conversational architecture shaping student engagement. More broadly, this work demonstrates how integrating generative AI annotation with advanced statistical modeling can scale analyses of tutoring dialogue and identify conversational practices that may prompt sustained engagement.
Kirk Vanacore, Jinsook Lee, Bakhtawar Ahtisham, Sarah Shaw, Justin Reich, René F. Kizilcec
L@S6
2026 Shiksha Copilot: Teacher-AI Collaboration for Curating and Customizing Lesson Plans in Low-Resource Schools CSCW038
abstract
This study investigates Shiksha Copilot, an AI-assisted lesson planning tool deployed in government schools across Karnataka, India. The system combined LLMs and human expertise through a structured process in which English and Kannada lesson plans were co-created by curators and AI; teachers then further customized these curated plans for their classrooms using their own expertise alongside AI support. Drawing on a large-scale mixed-methods study involving 1,043 teachers and 23 curators, we examine how educators collaborate with AI to generate context-sensitive lesson plans, assess the quality of AI-generated content, and analyze shifts in teaching practices within multilingual, low-resource environments. Our findings show that teachers used Shiksha Copilot both to meet administrative documentation needs and to support their teaching. The tool eased bureaucratic workload, reduced lesson planning time, and lowered teaching-related stress, while promoting a shift toward activity-based pedagogy. However, systemic challenges such as staffing shortages and administrative demands constrained broader pedagogical change. We frame these findings through the lenses of teacher-AI collaboration and communities of practice to examine the effective integration of AI tools in teaching. Finally, we propose design directions for future teacher-centered EdTech, particularly in multilingual and Global South contexts.
Deepak Varuvel Dennison, Bakhtawar Ahtisham, Kavyansh Chourasia, Nirmit Arora, René F. Kizilcec, Akshay Uttama Nambi, Tanuja Ganu, Aditya Vashistha
Proc. ACM Hum. Comput. Interact.6
2025 Benchmarking Generative AI for Scoring Medical Student Interviews in Objective Structured Clinical Examinations (OSCEs)
Jadon Geathers, Yann Hicke, Colleen E. Chan, Niroop Rajashekar, Sarah Young, Justin Sewell, Susannah Cornes, René F. Kizilcec, Dennis L. Shung
AIED (3)8
2025 Understanding Student Engagement with Large Language Model-Powered Course Assistants
Chang Liu 0122, Loc Hoang, Andrew Stolman, René F. Kizilcec, Bo Wu 0002
AIED (6)4
2025 Evaluating an AI Tutor for Bias Across Different Foundation Models
Aditya Vinodh, Emma Harvey, Husni Almoubayyed, Renzhe Yu, Christopher Brooks 0001, Allison Koenecke, René F. Kizilcec
AIED (6)7
2025 "Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education
Emma Harvey, Allison Koenecke, René F. Kizilcec
CHI3
2025 Applying DebiasEd: A Package for Mitigating Unfairness in Educational Data
Jade Cock, Frank Stinar, René F. Kizilcec, Tanja Käser
EDM3
2025 Understanding Predictive Models of Student Success with a Multiverse Analysis
Yunxuan Tang, Emma Harvey, Chengyuan Yao, Renzhe Yu, René F. Kizilcec, Christopher Brooks 0001
EDM5
2025 Algorithm Appreciation in Education: Educators Prefer Complex over Simple Algorithms
Kimberly Williamson, René F. Kizilcec, Sean Fath, Neil T. Heffernan
LAK2
2025 Fairness Over Time: A Nationwide Study of Evolving Bias in Dropout Prediction
abstract
The use of student learning data to predict educational outcomes has been widely studied, both in terms of model performance and fairness. An example of these predictive models is the use of Early Warning Systems (EWS), which identify students at risk of dropping out. They can be used continuously to make predictions, from the time of first enrollment until years into a degree program, to provide timely support. However, changes to student composition and their learning trajectories can alter the performance and group fairness of predictions over time. Using a nationwide higher education dataset, we examine changes in the fairness of a dropout prediction model at various points along the academic calendar. Our findings reveal that fairness is not static but evolves over time: the largest differences in AUC occur at 12 months after enrollment, a common evaluation point for dropout EWS. We discuss implications for the continued assessment of fairness in predictive algorithms in education.
Tereza Blazkova, René F. Kizilcec, Magnus Lindgaard Nielsen, David Dreyer Lassen, Andreas Bjerre-Nielsen
L@S2
2025 ChitterChatter: Curriculum-Aligned AI Speaking Partners for Language Learning Classrooms
abstract
Despite the importance of speaking practice in language learning, most students struggle to find low-stakes opportunities for authentic oral communication. ChitterChatter addresses this challenge by providing an AI-powered tool that enables instructors to create curriculum-aligned, voice-enabled conversation activities for students. Built on OpenAI's Realtime API and designed through iterative feedback from language education experts, ChitterChatter offers personalized, adaptive speaking practice while maintaining a judgment-free environment that promotes student comfort and confidence. Our pilot study with university-level Spanish learners shows that students value the platform's ability to provide authentic conversation practice without fear of judgment, although barriers to adoption remain. This paper presents ChitterChatter's design, preliminary evaluation results, and future directions for enhancing the system. Our findings demonstrate the potential of AI conversation partners to support classroom language instruction by increasing both the quantity and quality of speaking practice opportunities for students.
Jadon Geathers, A. J. Alvero, René F. Kizilcec
L@S3
2025 What Medical Students Need from Simulation: Insights to Guide Scalable Learning Design
abstract
Simulation-based learning (SBL) is a foundational component of clinical education, yet its implementation often varies in authenticity and educational value. Through semi-structured interviews with ten medical students across three U.S. institutions, we examined how students engage with SBL within their broader learning contexts. Our thematic analysis identified adaptive learning strategies developed in response to time constraints, limited formal guidance, and a fragmented educational landscape. Students described challenges including gaps in simulation realism, inconsistent assessment objectives, and difficulty obtaining actionable feedback. This study provides critical learner-centered design insights intended to inform the development of scalable solutions-particularly digital or AI-driven platforms-that can address these limitations and better support learning in high-pressure professional education.
Jadon Geathers, Yann Hicke, Naphasjutha Kongsonthana, Justin Sewell, Anyanate Gwendolyne Jack, Dennis L. Shung, MacKenzi Preston, Susannah Cornes, René F. Kizilcec
L@S9
2025 Investigating Systematic Variation in Academic Procrastination Behavior by Course, Assignment, and Student Characteristics
abstract
Procrastination has been linked to lower academic performance and sociodemographic achievement gaps in a variety of educational contexts, posing challenges to student success and educational equity. While prior research acknowledges that learning environments play a crucial role in shaping student procrastination alongside personal traits, there is a lack of solid empirical evidence on the connection between specific variations in learning environments and academic procrastination. This study provides a large-scale evaluation of the relationship between course and assignment characteristics and student procrastination behavior using a sample of 33,514 students across 3,169 courses at a US university. Using fixed effects linear regression models, we find that students tend to procrastinate less in courses with larger enrollment, non-introductory content, and well-structured deadlines. Procrastination is also lower for assignments with spaced-out deadlines, weekend deadlines, and a quiz or discussion post format. However, these patterns do not apply equally across all student groups. Male, ethnic minority, and first-generation college students exhibit higher levels of procrastination than their peers, especially for courses and assignments with specific characteristics. We suggest two instructional design strategies to help manage procrastination across student populations: (1) allowing more time before the first assignment deadline, and (2) ensuring adequate spacing between deadlines. This study provides large-scale evidence of the complex relationship between learning environment design, student characteristics, and procrastination.
Yan Tao, Nathan Maidi, Renzhe Yu, René F. Kizilcec
L@S4
2025 Advancing the Science of Teaching with Tutoring Data: A Collaborative Workshop with the National Tutoring Observatory
abstract
L@S ’25, Palermo, Italy
Danielle R. Thomas, Dorottya Demszky, Kenneth R. Koedinger, Josh Marland, Doug Pietrzak, Justin Reich, Rachel Slama, Amalia Christina Toutziaridi, René F. Kizilcec
L@S9
2024 Human-Algorithmic Interaction Using a Large Language Model-Augmented Artificial Intelligence Clinical Decision Support System
abstract
Integration of artificial intelligence (AI) into clinical decision support systems (CDSS) poses a socio-technological challenge that is impacted by usability, trust, and human-computer interaction (HCI). AI-CDSS interventions have shown limited benefit in clinical outcomes, which may be due to insufficient understanding of how health-care providers interact with AI systems. Large language models (LLMs) have the potential to enhance AI-CDSS, but haven’t been studied in either simulated or real-world clinical scenarios. We present findings from a randomized controlled trial deploying AI-CDSS for the management of upper gastrointestinal bleeding (UGIB) with and without an LLM interface within realistic clinical simulations for physician and medical student participants. We find evidence that LLM augmentation improves ease-of-use, that LLM-generated responses with citations improve trust, and HCI varies based on clinical expertise. Qualitative themes from interviews suggest the perception of LLM-augmented AI-CDSS as a team-member used to confirm initial clinical intuitions and help evaluate borderline decisions.
Niroop Rajashekar, Yeo Eun Shin, Yuan Pu 0002, Sunny Chung, Kisung You, Mauro Giuffrè, Colleen E. Chan, Theo Saarinen, Allen Hsiao, Jasjeet S. Sekhon, Ambrose Wong, Leigh V. Evans, René F. Kizilcec, Loren Laine, Terika McCall, Dennis L. Shung
CHI13
2024 Which Planning Tactics Predict Online Course Completion?
abstract
Planning is a self-regulated learning strategy and widely used behavior change technique that can help learners achieve academic goals (e.g., pass an exam, apply to college, or complete an online course). Numerous studies have tested the effects of planning interventions, but few have examined the content of learners’ plans and how it relates to their academic outcomes. Building on a large-scale intervention study, we conducted a qualitative content analysis of 650 learner plans sampled from 15 massive open online courses (MOOCs). We identified a number of planning tactics, compared their prevalence, and examined which ones significantly predict course progress and completion using regression analyses. We found that learners whose plans specify a time of day (e.g., morning, afternoon, night) are significantly more likely to complete a MOOC, but only 25% of the learners in our sample used this tactic. The high degree of variation in the effectiveness of planning tactics may contribute to mixed intervention findings in scale-up studies. Models of plan effectiveness can be used to provide feedback on the quality of learners’ plans and encourage them to use effective tactics to achieve their learning goals.
Ji Yong Cho, Yan Tao, Michael Yeomans, Dustin Tingley, René F. Kizilcec
LAK5
2023 The Role of Gender in Students' Privacy Concerns about Learning Analytics: Evidence from five countries
abstract
The protection of students’ privacy in learning analytics (LA) applications is critical for cultivating trust and effective implementations of LA in educational environments around the world. However, students’ privacy concerns and how they may vary along demographic dimensions that historically influence these concerns have yet to be studied in higher education. Gender differences, in particular, are known to be associated with people's information privacy concerns, including in educational settings. Building on an empirically validated model and survey instrument for student privacy concerns, their antecedents and their behavioral outcomes, we investigate the presence of gender differences in students’ privacy concerns about LA. We conducted a survey study of students in higher education across five countries (N = 762): Germany, South Korea, Spain, Sweden and the United States. Using multiple regression analysis, across all five countries, we find that female students have stronger trusting beliefs and they are more inclined to engage in self-disclosure behaviors compared to male students. However, at the country level, these gender differences are significant only in the German sample, for Bachelor's degree students, and for students between the ages of 18 and 24. Thus, national context, degree program, and age are important moderating factors for gender differences in student privacy concerns.
René F. Kizilcec, Olga Viberg, Ioana Jivet, Alejandra Martínez-Monés, Alice Oh, Stefan Hrastinski, Chantal Mutimukwe, Maren Scheffel
LAK1
2023 Fourth Annual Workshop on A/B Testing and Platform-Enabled Learning Research
Steven Ritter 0001, Neil T. Heffernan, Joseph Jay Williams, Derek Lomas, Klinton Bicknell, Jeremy Roschelle, Benjamin Motz 0002, Danielle S. McNamara, Richard G. Baraniuk, Debshila Basu Mallick, René F. Kizilcec, Ryan Baker 0001, Stephen Fancsali, April Murphy
L@S11
2023 Developing a Prototype to Scale up Digital Support for Online Assessment Design
abstract
Educators rarely have access to just-in-time feedback and guiding heuristics when designing or updating assessments in higher education. This study describes the initial development process for an automated support system for designing high-quality online assessments. We identify key elements to embed in this digital artifact to offer just-in-time support for educators to design and evaluate their online assessments. We follow a design science approach in six stages, because it simultaneously generates knowledge about the method used to develop the artifact and the design of the artifact itself. Specifically, we focus on the early stages of problem identification, solution objectives, and initial conceptual design. After reviewing multiple assessment models and frameworks, we discuss a recent framework for evaluating and designing high-quality online assessments, consisting of ten design and contextual elements. This framework underpins the proposed solution which is a digital artifact that encourages consideration of alternate forms of assessment while retaining the flexibility to operate within individual educators' design practices and contexts. We expect the proposed system to help educators and instructional designers to better understand the strengths and weaknesses of their assessments, consider alternate forms of assessment, and incorporate the system into their assessment design process.
Andrew Cram, Corina Raduescu, Sandris Zeivots, Adele Smolansky, René F. Kizilcec, Elaine Huber
L@S5
2023 Evaluating a Learned Admission-Prediction Model as a Replacement for Standardized Tests in College Admissions
abstract
A growing number of college applications has presented an annual challenge for college admissions in the United States. Admission offices have historically relied on standardized test scores to organize large applicant pools into viable subsets for review. However, this approach may be subject to bias in test scores and selection bias in test-taking with recent trends toward test-optional admission. We explore a machine learning-based approach to replace the role of standardized tests in subset generation while taking into account a wide range of factors extracted from student applications to support a more holistic review. We evaluate the approach on data from an undergraduate admission office at a selective US institution (13,248 applications). We find that a prediction model trained on past admission data outperforms an SAT-based heuristic and matches the demographic composition of the last admitted class. We discuss the risks and opportunities for how such a learned model could be leveraged to support human decision-making in college admissions.
René F. Kizilcec, Thorsten Joachims
L@S2
2023 Educator and Student Perspectives on the Impact of Generative AI on Assessments in Higher Education
abstract
The sudden popularity and availability of generative AI tools, such as ChatGPT that can write compelling essays on any topic, code in various programming languages, and ace standardized tests across domains, raises questions about the sustainability of traditional assessment practices. To seize this opportunity for innovation in assessment practice, we conducted a survey to understand both the educators' and students' perspectives on the issue. We measure and compare attitudes of both stakeholders across various assessment scenarios, building on an established framework for examining the quality of online assessments along six dimensions. Responses from 389 students and 36 educators across two universities indicate moderate usage of generative AI, consensus for which types of assessments are most impacted, and concerns about academic integrity. Educators prefer adapted assessments that assume AI will be used and encourage critical thinking, but students' reaction is mixed, in part due to concerns about a loss of creativity. The findings show the importance of engaging educators and students in assessment reform efforts to focus on the process of learning over its outputs, higher-order thinking, and authentic applications.
Adele Smolansky, Andrew Cram, Corina Raduescu, Sandris Zeivots, Elaine Huber, René F. Kizilcec
L@S6
2022 A Review of Learning Analytics Dashboard Research in Higher Education: Implications for Justice, Equity, Diversity, and Inclusion
abstract
Learning analytics dashboards (LADs) are becoming more prevalent in higher education to help students, faculty, and staff make data-informed decisions. Despite extensive research on the design and implementation of LADs, few studies have investigated their relation to justice, equity, diversity, and inclusion (JEDI). Excluding these issues in LAD research limits the potential benefits of LADs generally and risks reinforcing long-standing inequities in education. We conducted a critical literature review, identifying 45 relevant papers to answer three research questions: how is LAD research improving JEDI, ii. how might it maintain or exacerbate inequitable outcomes, and iii. what opportunities exist in this space to improve JEDI in higher education. Using thematic analysis, we identified four common themes: (1) participant identities and researcher positionality, (2) surveillance concerns, (3) implicit pedagogies, and (4) software development resources. While we found very few studies directly addressing or mentioning JEDI concepts, we used these themes to explore ways researchers could consider JEDI in their studies. Our investigation highlights several opportunities to intentionally incorporate JEDI into LAD research by sharing software resources and conducting cross-border collaborations, better incorporating user needs, and centering considerations of justice in LAD efforts to improve historical inequities.
Kimberly Williamson, René F. Kizilcec
LAK2
2022 Pathways: Exploring Academic Interests with Historical Course Enrollment Records
abstract
Students are encouraged to explore their interests during college to stimulate intellectual growth and prepare for a dynamic labor market. However, interest exploration is entangled with the fateful process of choosing courses for enrollment, and most institutions offer limited tools to help students choose. We propose Pathways, an interactive course information retrieval tool that facilitates interest exploration and course discovery with a diverse pool of historical course enrollment records. The tool visualizes sequences of course enrollments as "academic pathways" to grant students unprecedented insights into the academic choices of prior students. We share our design process, including a formative study on need analysis, the UX and algorithm design, and an evaluation study. We find that Pathways supports students in finding courses that both match their interests and expose them to new ideas. We discuss directions for future work on how interest exploration can be promoted at scale and on how to utilize historical course enrollment data through visualization.
Youjie Chen, Annie Fu, Jennifer Jia-Ling Lee, Ian Wilkie Tomasik, René F. Kizilcec
L@S5
2022 Measuring Cultural Dimensions of Learning in Online Courses
abstract
Online courses lower geographic barriers to educational access and attract learners from around the world. The resulting cultural diversity in online courses has implications for learning preferences, behaviors and outcomes, but established measures of culture are not adapted to educational contexts. We adapted and tested a survey instrument of cultural dimensions of learning that is grounded in cultural psychology research and spans four dimensions: knowledge construction, pedagogical orientation, uncertainty tolerance, and consensus building. We collected 600 responses in two online courses, conducted an explanatory factor analysis, and compared responses across five countries. We found that the instrument has a clear factor structure with high internal consistency, and it can distinguish cultures between countries. The instrument can be used to better understand learners and their culture in the process of course design and evaluation.
Ji Yong Cho, Yue Li 0051, Marianne E. Krasny, René F. Kizilcec
L@S4
2022 Effects of Framing Professional Development as a Career Growth Opportunity on Course Completion
abstract
Professional development (PD) trainings help ensure employees keep up with important changes in practice, policy, and technology, but they are often perceived as burdensome by employees, likely contributing to compliance issues. Negative attitudes towards PD trainings may arise because employees view them as a chore rather than a benefit. We conducted a multi-faceted utility-value intervention in the context of a mandated, state-wide training program over two years. The intervention encouraged participants to see PD training as an opportunity for professional growth using messages embedded in email and on the PD website. We randomly assigned 98 employers (496 employees) to either the intervention condition or a business-as-usual control condition. We found limited evidence of the intervention increasing course completion. Qualitative findings suggest alternative interventions to address time management and structural barriers in trainings and workplaces.
René F. Kizilcec, Jennifer A. Mimno, Andrew J. Karhan
L@S1
2022 Third Annual Workshop on A/B Testing and Platform-Enabled Learning Research
abstract
Learning engineering adds tools and processes to learning platforms to support improvement research. One kind of tool is A/B testing, which is common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE). A number of A/B testing systems focused on educational applications have arisen recently, including UpGrade and E-TRIALS. A/B testing can be part of the puzzle of how to improve educational platforms, and yet challenging issues in education go beyond the generic paradigm. For example, the importance of teachers and instructors to learning means that students are not only connecting with software as individuals, but also as part of a shared classroom experience. Further, learning in topics like mathematics can be highly dependent on prior learning, and thus A or B may not be better overall, but only in interaction with prior knowledge. In response, a set of learning platforms is opening their systems to improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore how A/B testing in educational contexts is different, how learning platforms are opening up new possibilities, and how these empirical approaches can be used to drive powerful gains in student learning. It will also discuss forthcoming opportunities for funding to conduct platform-enabled learning research.
Steven Ritter 0001, Neil T. Heffernan, Joseph Jay Williams, Derek Lomas, Benjamin Motz 0002, Debshila Basu Mallick, Klinton Bicknell, Danielle S. McNamara, René F. Kizilcec, Jeremy Roschelle, Richard G. Baraniuk, Ryan Baker 0001
L@S9
2022 Large-Scale Student Data Reveal Sociodemographic Gaps in Procrastination Behavior
abstract
University students have to manage complex and demanding schedules to keep up with coursework across multiple classes while navigating formative personal, cultural, and financial events. Procrastination, the act of deferring study effort until the task deadline, is therefore a prevalent phenomenon, but whether it is more common among historically disadvantaged students is unknown. If systematic differences in procrastination behavior exist across sociodemographic groups, they may also contribute to achievement gaps, considering that procrastination is largely negatively associated with academic performance in prior research. We therefore investigate these questions in the context of assignment submission using campus-wide learning management system (LMS) data from a large U.S. research university. We analyze 2,631,893 submission records by 25,659 students across 2,153 courses and propose a context-agnostic procrastination score for each student in each course based on their assignment submission times relative to classmates. Based on this procrastination score, we find significantly higher levels of procrastination behavior among males, racial minorities, and first-generation college students than their peers. However, these differences only explain performance gaps to a very limited extent and the negative association between procrastination behavior and performance remains relatively stable across student groups. This large-scale behavioral study advances the understanding of academic procrastination through an equity lens and informs the development of scalable interventions to mitigate the negative effects of procrastination.
Sunil Sabnis, Renzhe Yu, René F. Kizilcec
L@S3
2022 Large-scale Analysis of Discussion Networks in College Courses
abstract
Online discussion boards serve an important role in college courses by facilitating social learning and student support. However, the student experience and learning outcomes are likely to depend on the structure of student engagement with one another and the teaching staff. Using social network analysis, we investigated the network structure of 616 course discussion boards at a selective research university. We first examine variation in discussion boards using a wide range of composite metrics from the social network analysis literature. We then develop a typology of discussion board networks using principal component analysis and k-Means clustering to arrive at three clusters: dense discussion; distinct discussion groups; and discussion brokers and hubs.
Kimberly Williamson, René F. Kizilcec
L@S2
2021 Effects of Algorithmic Transparency in Bayesian Knowledge Tracing on Trust and Perceived Accuracy
Kimberly Williamson, René F. Kizilcec
EDM2
2021 Applying the Behavior Change Technique Taxonomy from Public Health Interventions to Educational Research
abstract
Public health research has developed a deep understanding of ways to help people live healthier lives through scalable interventions that change their behaviors. This work offers valuable insights for supporting learners in educational contexts, especially for improving self-regulation and goal-directed behaviors like completing a course of study--a persistent issue in formal and information post-secondary education. We present the widely adopted Behavior Change Technique (BCT) taxonomy as a model for systematically cataloging interventions in education and as a resource for inspiring new interventions in education based on public health evidence. Approaching the issue of learner attrition from the BCT perspective, we show how recent educational interventions fit into the BCT taxonomy and how the taxonomy can be used to develop new evidence-based intervention approaches. Borrowing insights from decades of public health research can advance parallel efforts in education to help learners at scale to stay on track and reach their academic goals.
Ji Yong Cho, René F. Kizilcec
L@S2
2021 Using Social Norms to Promote Actions Beyond the Course
abstract
Educators and researchers in online education have grappled with not only how to increase course completion but also how to make a broader impact that goes beyond online courses, such as course participants' real-world applications of the learned knowledge and skills. Research in social psychology and behavioral science suggests that social norms interventions, which convey norms shared in the community that people belong in to promote desirable behaviors, can offer a low-cost and scalable approach to encourage actions beyond the courses (ABCs). We tested three social norm interventions that presented a weekly normative message (descriptive, dynamic, or injunctive norm) with aggregate information about course participants' ABCs in the prior week. Randomized experiments in three online courses found effects on ABCs to be weak and moderated by norm message type and the complexity of the target behavior. Although the interventions did not improve course completion, the dynamic norm message was more effective at promoting ABCs for complex behaviors, such as developing environmental education activities.
Ji Yong Cho, Yue Li 0051, Anne K. Armstrong, Alex Russ, Marianne E. Krasny, René F. Kizilcec
L@S6
2021 Student Perceptions of Social Support in the Transition to Emergency Remote Instruction
abstract
University courses around the world suddenly transitioned to emergency remote instruction in response to the COVID-19 pandemic. We study changes in students' experience of support from their instructors and peers in large lecture courses. Social support can act as an important resource for students and buffer against mental distress. We find that students experienced more support from instructors but less support from their peers after the transition to remote instruction. Remote learning was less active and involved fewer peer interactions, with synchronous classes resembling online office hours and students struggling to get help. Our findings suggest the need for additional resources to help students stay connected and facilitate collaboration online.
Ji Yong Cho, Ian Wilkie Tomasik, René F. Kizilcec
L@S4
2021 Learning Analytics Dashboard Research Has Neglected Diversity, Equity and Inclusion
abstract
Learning analytic dashboards (LADs) have become more prevalent in higher education to help students, faculty, and staff make data-informed decisions. Despite extensive research on the design and usability of LADs, few studies have examined them in relation to issues of diversity, equity, and inclusion. We conducted a critical literature review to address three research questions: How does LAD research contribute to improving diversity, equity, and inclusion? How might LADs contribute to maintaining or exacerbating inequitable outcomes? And what future opportunities exist in this research space? Our review showed little use of LADs to address or improve issues of diversity, equity, and inclusion in the literature thus far. We argue that excluding these issues from LAD research is not an isolated oversight and it risks reinforcing existing inequities within the higher education system. We argue that LADs can be designed, researched, and deployed intentionally to advance equitable outcomes and help dismantle inequities in education. We highlight opportunities for future LAD research to address issues of diversity, equity, and inclusion.
Kimberly Williamson, René F. Kizilcec
L@S2
2021 Should College Dropout Prediction Models Include Protected Attributes?
abstract
Early identification of college dropouts can provide tremendous value for improving student success and institutional effectiveness, and predictive analytics are increasingly used for this purpose. However, ethical concerns have emerged about whether including protected attributes in these prediction models discriminates against underrepresented student groups and exacerbates existing inequities. We examine this issue in the context of a large U.S. research university with both residential and fully online degree-seeking students. Based on comprehensive institutional records for the entire student population across multiple years (N = 93,457), we build machine learning models to predict student dropout after one academic year of study and compare the overall performance and fairness of model predictions with or without four protected attributes (gender, URM, first-generation student, and high financial need). We find that including protected attributes does not impact the overall prediction performance and it only marginally improves the algorithmic fairness of predictions. These findings suggest that including protected attributes is preferable. We offer guidance on how to evaluate the impact of including protected attributes in a local context, where institutional stakeholders seek to leverage predictive analytics to support student success.
Renzhe Yu, René F. Kizilcec
L@S3
2021 Investigating Technostress Among Teachers in Low-Income Indian Schools
abstract
Smartphones play an increasingly large role in the professional lives of teachers in low-income contexts, creating an urgent need to better understand the role of technology-related stress (technostress) in teachers' smartphone use for work. We contribute a mixed methods study analyzing the impact of smartphone use on teachers' work lives in low-income Indian schools. Findings from 70 interviews and 1,361 survey responses suggest that although smartphones aid teaching and administrative functions, smartphone use also significantly predicts burnout among teachers, with technostress providing a major explanation for this relationship. We reveal how teachers' work is constantly surveilled and monitored via technology and how teachers' personal smartphones were controlled and repurposed through socio-technical structures by the higher management to serve management's goals, substantially increasing the work teachers were required to perform outside of work hours. Our work extends technostress research to HCI4D contexts and highlights the need to develop better support structures for teachers and rethink how smartphones are used in their work.
Rama Adithya Varanasi, Aditya Vashistha, René F. Kizilcec, Nicola Dell
Proc. ACM Hum. Comput. Interact.3
2020 Welcome to the Course: Early Social Cues Influence Women's Persistence in Computer Science
abstract
First impressions influence subsequent behavior, especially when deciding how much effort to invest in an activity such as taking an online course. In computer programming courses, a context where social group stereotypes are salient, social cues early in the course can be used strategically to affirm members of historically underrepresented groups in their sense of belonging. We tested this idea in two randomized field experiments (N=53,922) by varying the social identity and status of the presenter of a welcome video and assessing online learners' persistence and achievement. Counter to our hypotheses, we found lower persistence among women in certain age groups if the welcome video was presented by a female instructor or by lower-status peers. Men remained unaffected. The results suggest that women are more responsive to social cues in online STEM courses, an environment where their social identity has been negatively stereotyped. Presenting a male and female instructor together was an effective strategy for retaining women in the course.
René F. Kizilcec, Andrew J. Saltarelli, Petra Bonfert-Taylor, Michael Goudzwaard, Ella Hamonic, Rémi Sharrock
CHI1
2020 Return of the Student: Predicting Re-Engagement in Mobile Learning
Maximillian Chen 0002, René F. Kizilcec
EDM2
2020 Designing Inclusive Learning Environments
abstract
Large-scale online learning environments present new opportunities to address the need for greater inclusivity in education. Unlike residential environments, which have physical and logistic constraints (e.g., classroom configurations, sizes, and scheduling) that impede our ability to enact more inclusive pedagogy, online learning environments can be personalized and adapted to individual learner needs. As these environments are completely technology mediated, they offer an almost infinite design space for innovation. Social-scientific research on inclusivity in residential settings provides insight into how we might design for online learning environments, however evidence of efficacious digital implementations of these insights is limited. This workshop aims to advance our understanding of the ways in which adaptivity can be leveraged to buttress inclusivity in STEM learning. Through brief paper presentations and collaborative activities we intend to outline design opportunities in the scaled learning space for creating more inclusive environments.
Christopher Brooks 0001, René F. Kizilcec, Nia Nixon
L@S2
2020 Examining Sources of Variation in Student Confusion in College Classes
abstract
Students often experience confusion while learning, and if promptly resolved, it can promote engagement and deeper understanding. However, detecting student confusion and intervening in a timely and scalable manner challenges even seasoned instructors. To understand when and where students are most likely to be confused, we study the systematic occurrence of confusion in college classes among 29,511 students in twelve universities. We use a novel method for affect detection that allows students to self-report confusion on individual presentation slides during their classes. Across 1,366 class presentations, we find that confusion arises at different times during class and depends on class duration, class size, type of institution, and academic discipline. Confusion is most prevalent during short presentations, in small classes, low-tier institutions, and scientific disciplines.
Youjie Chen, René F. Kizilcec
L@S2
2020 Student Engagement in Mobile Learning via Text Message
abstract
Mobile learning is expanding rapidly due to its accessibility and affordability, especially in resource-poor parts of the world. Yet how students engage and learn with mobile learning has not been systematically analyzed at scale. This study examines how 93,819 Kenyan students in grades 6, 9, and 12 use a text message-based mobile learning platform that has millions of users across Sub-Saharan Africa. We investigate longitudinal variation in engagement over a one-year period for students in different age groups and check for evidence of learning gains using learning curve analysis. Student engagement is highest during school holidays and leading up to standardized exams, but persistence over time is low: under 25% of students return to the platform after joining. Clustering students into three groups based on their level of activity, we examine variation in their learning behaviors and quiz performance over their first ten days. Highly active students exhibit promising trends in terms of quiz completion, reattempts, and accuracy, but we do not see evidence of learning gains in this study. The findings suggest that students in Kenya use mobile learning either as an ad-hoc resource or a low-cost tutor to complement formal schooling and bridge gaps in instruction.
René F. Kizilcec, Maximillian Chen 0002
L@S1
2019 Psychologically Inclusive Design: Cues Impact Women's Participation in STEM Education
abstract
Visual and verbal cues can reinforce barriers to access for women in science, technology, engineering, and math (STEM) disciplines. Psychologically inclusive design is an evidence-based approach to reduce psychological barriers by strategically placing content and design cues in the environment. Two large field experiments provide estimates of the behavioral impact of psychologically inclusive cues on women's and men's enrollment behaviors in an online learning environment. First, a gender-inclusive photo and statement in an online advertisement for a STEM course increased the click-through rate among women but not men by 26% (N=209,000). Second, an inclusivity statement with a gender-inclusive course image to the enrollment page raised the proportion of women enrolling in a STEM course by up to 18% (N=63,000). These findings contribute evidence of the behavioral impact of psychologically inclusive design to the literature and yield practical implications for the presentation of STEM opportunities.
René F. Kizilcec, Andrew J. Saltarelli
CHI1
2019 Growth Mindset Predicts Student Achievement and Behavior in Mobile Learning
abstract
Students' personal qualities other than cognitive ability are known to influence persistence and achievement in formal learning environments, but the extent of their influence in digital learning environments is unclear. This research investigates non-cognitive factors in mobile learning in a resource-poor context. We surveyed 1,000 Kenyan high school students who use a popular SMS-based learning platform that provides formative assessments aligned with the national curriculum. Combining survey responses with platform interaction logs, we find growth mindset to be one of the strongest predictors of assessment scores. We investigate theory-based behavioral mechanisms to explain this relationship. Although students who hold a growth mindset are not more likely to persist after facing adversity, they spend more time on each assessment, increasing their likelihood of answering correctly. Results suggest that cultivating a growth mindset can motivate students in a resource-poor context to excel in a mobile learning environment.
René F. Kizilcec, Daniel Goldfarb
L@S1
2019 Can a diversity statement increase diversity in MOOCs?
abstract
Despite the fact that anyone can sign up for open online courses, their enrollment patterns reflect the historical underrepresentation of certain sociodemographic groups (e.g. women in STEM disciplines). We theorize that enrollment choices online are shaped by contextual cues that activate stereotypes about numeric representation and climate in brick-and-mortar institutions. A longitudinal matched-pairs experiment with 14 MOOCs (N=29,000) tested this theory by manipulating the presence of a diversity statement on course pages and measuring effects on who enrolls. We found a 3% increase in the proportion of students with lower socioeconomic status. The effect size varied across courses between -0.5 and 7 percentage points. No significant changes in enrollment patterns by gender, age, and national development level occurred. Implications for the use and content of diversity statements and their alternatives are discussed.
René F. Kizilcec, Andrew J. Saltarelli
L@S1
2019 How Teachers in India Reconfigure their Work Practices around a Teacher-Oriented Technology Intervention
abstract
The proliferation of mobile devices around the world, combined with falling costs of hardware and Internet connectivity, have resulted in an increasing number of organizations that work to introduce educational technology interventions into low-income schools in the Global South. However, to date, most prior HCI research examining such interventions has focused on interventions that target students. In this paper, we expand prior literature by examining an intervention, called Meghshala, that targets teachers in low-income schools as its primary users. Through interviews and observations with 39 participants from 12 government schools in India, we show how the introduction of a teacher-focused technology intervention causes teachers to reconfigure their work practices, including lesson preparation, in-classroom teaching practices, bureaucratic work processes, and post-teaching feedback mechanisms. We use the concept of material agency to analyze our findings with respect to teacher agency and reconfiguration, and use theories of teacher knowledge to highlight the kinds of knowledge production that teachers in our research context tend to focus on (e.g., content knowledge). Finally, we offer design opportunities for future teacher-focused technology interventions.
Rama Adithya Varanasi, René F. Kizilcec, Nicola Dell
Proc. ACM Hum. Comput. Interact.2
2018 Social Influence and Reciprocity in Online Gift Giving
abstract
Giving gifts is a fundamental part of human relationships that is being affected by technology. The Internet enables people to give at the last minute and over long distances, and to observe friends giving and receiving gifts. How online gift giving spreads in social networks is therefore important to understand. We examine 1.5 million gift exchanges on Facebook and show that receiving a gift causes individuals to be 56% more likely to give a gift in the future. Additional surveys show that online gift giving was more socially acceptable to those who learned about it by observing friends' participation instead of a non-social encouragement. Most receivers pay the gift forward instead of reciprocating directly online, although surveys revealed additional instances of direct reciprocity, where the initial gifting occurred offline. Thus, social influence promotes the spread of online gifting, which both complements and substitutes for offline gifting.
René F. Kizilcec, Eytan Bakshy, Dean Eckles, Moira Burke
CHI1
2018 The half-life of MOOC knowledge: a randomized trial evaluating knowledge retention and retrieval practice in MOOCs
abstract
Retrieval practice has been established in the learning sciences as one of the most effective strategies to facilitate robust learning in traditional classroom contexts. The cognitive theory underpinning the "testing effect" states that actively recalling information is more effective than passively revisiting materials for storing information in long-term memory. We document the design, deployment, and evaluation of an Adaptive Retrieval Practice System (ARPS) in a MOOC. This push-based system leverages the testing effect to promote learner engagement and achievement by intelligently delivering quiz questions from prior course units to learners throughout the course. We conducted an experiment in which learners were randomized to receive ARPS in a MOOC to track their performance and behavior compared to a control group. In contrast to prior literature, we find no significant effect of retrieval practice in this MOOC environment. In the treatment condition, passing learners engaged more with ARPS but exhibited similar levels of knowledge retention as non-passing learners.
Dan Davis, René F. Kizilcec, Claudia Hauff, Geert-Jan Houben
LAK2
2018 How a data-driven course planning tool affects college students' GPA: evidence from two field experiments
abstract
College students rely on increasingly data-rich environments when making learning-relevant decisions about the courses they take and their expected time commitments. However, we know little about how their exposure to such data may influence student course choice, effort regulation, and performance. We conducted a large-scale field experiment in which all the undergraduates at a large, selective university were randomized to an encouragement to use a course-planning web application that integrates information from official transcripts from the past fifteen years with detailed end-of-course evaluation surveys. We found that use of the platform lowered students' GPA by 0.28 standard deviations on average. In a subsequent field experiment, we varied access to information about course grades and time commitment on the platform and found that access to grade information in particular lowered students' overall GPA. Our exploratory analysis suggests these effects are not due to changes in the portfolio of courses that students choose, but rather by changes to their behavior within courses.
Sorathan Chaturapruek, Thomas S. Dee, Ramesh Johari, René F. Kizilcec, Mitchell L. Stevens
L@S4
2017 Follow the successful crowd: raising MOOC completion rates through social comparison at scale
abstract
Social comparison theory asserts that we establish our social and personal worth by comparing ourselves to others. In in-person learning environments, social comparison offers students critical feedback on how to behave and be successful. By contrast, online learning environments afford fewer social cues to facilitate social comparison. Can increased availability of such cues promote effective self-regulatory behavior and achievement in Massive Open Online Courses (MOOCs)? We developed a personalized feedback system that facilitates social comparison with previously successful learners based on an interactive visualization of multiple behavioral indicators. Across four randomized controlled trials in MOOCs (overall N = 33, 726), we find: (1) the availability of social comparison cues significantly increases completion rates, (2) this type of feedback benefits highly educated learners, and (3) learners' cultural context plays a significant role in their course engagement and achievement.
Dan Davis, Ioana Jivet, René F. Kizilcec, Guanliang Chen, Claudia Hauff, Geert-Jan Houben
LAK3
2017 Towards Equal Opportunities in MOOCs: Affirmation Reduces Gender & Social-Class Achievement Gaps in China
abstract
The presence of achievement gaps in Massive Open Online Courses (MOOCs) implies that not everyone who can gain access to a course shares the same opportunities to succeed. This study advances research on a social psychological barrier to achievement that exists alongside important structural barriers (e.g., Internet access, insufficient prior knowledge). Learners who experience social identity threat (SIT) - a fear of being judged negatively in light of a social group they identify with - are at risk of underperforming. An initial survey identified lower-class men as an at-risk group in an English language learning MOOC for Chinese learners (N = 1,664). In a subsequent randomized experiment, an interdependent value relevance affirmation intervention raised grades, persistence, and completion rates exclusively among lower-class men - the lowest performing group in the course (N = 1,990). Efforts to establish equal opportunities in online learning should go beyond initiatives that increase access through technology to incorporate strategies that lower psychological barriers to create safe and inclusive learning environments.
René F. Kizilcec, Glenn M. Davis, Geoffrey L. Cohen
L@S1
2016 How Much Information?: Effects of Transparency on Trust in an Algorithmic Interface
abstract
The rising prevalence of algorithmic interfaces, such as curated feeds in online news, raises new questions for designers, scholars, and critics of media. This work focuses on how transparent design of algorithmic interfaces can promote awareness and foster trust. A two-stage process of how transparency affects trust was hypothesized drawing on theories of information processing and procedural justice. In an online field experiment, three levels of system transparency were tested in the high-stakes context of peer assessment. Individuals whose expectations were violated (by receiving a lower grade than expected) trusted the system less, unless the grading algorithm was made more transparent through explanation. However, providing too much information eroded this trust. Attitudes of individuals whose expectations were met did not vary with transparency. Results are discussed in terms of a dual process model of attitude change and the depth of justification of perceived inconsistency. Designing for trust requires balanced interface transparency - not too little and not too much.
René F. Kizilcec
CHI1
2016 Recommending Self-Regulated Learning Strategies Does Not Improve Performance in a MOOC
abstract
Many committed learners struggle to achieve their goal of completing a Massive Open Online Course (MOOC). This work investigates self-regulated learning (SRL) in MOOCs and tests if encouraging the use of SRL strategies can improve course performance. We asked a group of 17 highly successful learners about their own strategies for how to succeed in a MOOC. Their responses were coded based on a SRL framework and synthesized into seven recommendations. In a randomized experiment, we evaluated the effect of providing those recommendations to learners in the same course (N = 653). Although most learners rated the study tips as very helpful, the intervention did not improve course persistence or achievement. Results suggest that a single SRL prompt at the beginning of the course provides insufficient support. Instead, embedding technological aids that adaptively support SRL throughout the course could better support learners in MOOCs.
René F. Kizilcec, Mar Pérez-Sanagustín, Jorge J. Maldonado
L@S1
2015 To Play or Not to Play: Interactions between Response Quality and Task Complexity in Games and Paid Crowdsourcing
abstract
Digital games are a viable alternative to accomplish crowdsourcing tasks that would traditionally require paid online labor. This study compares the quality of crowdsourcing with games and paid crowdsourcing for simple and complex annotation tasks in a controlled exper-iment. While no difference in quality was found for the simple task, paid contributors’ response quality was sub-stantially lower than players’ quality for the complex task (92% vs. 78% average accuracy). Results suggest that crowdsourcing with games provides similar and potentially even higher response quality relative to paid crowdsourcing.
Markus Krause, René F. Kizilcec
HCOMP2
2015 Attrition and Achievement Gaps in Online Learning
abstract
Attrition in online learning is generally higher than in traditional settings, especially in large-scale online learning environments. A systematic analysis of individual differences in attrition and performance in 20 massive open online courses (N > 67,000) revealed a geographic achievement gap and a gender achievement gap. Online learners in Africa, Asia, and Latin America scored substantially lower grades and were only half as likely to persist than those in Europe, Oceania, and Northern America. Women also exhibited lower persistence and performance than men. Yet more persistent learners were only marginally more satisfied with their achievement. The primary obstacle for most learners was finding time for the course, which was partly related to low levels of volitional control. Self-ascribed successful learners reported higher levels of goal striving, growth mindset, and feelings of social belonging than unsuccessful ones. Insights into why learners leave online courses inform models of attrition and targeted interventions to support learners achieve their goals.
René F. Kizilcec, Sherif A. Halawa
L@S1
2015 Motivation as a Lens to Understand Online Learners: Toward Data-Driven Design with the OLEI Scale
abstract
Open online learning environments attract an audience with diverse motivations who interact with structured courses in several ways. To systematically describe the motivations of these learners, we developed the Online Learning Enrollment Intentions (OLEI) scale, a 13-item questionnaire derived from open-ended responses to capture learners' authentic perspectives. Although motivations varied across courses, we found that each motivation predicted key behavioral outcomes for learners (N = 71, 475 across 14 courses). From learners' motivational and behavioral patterns, we infer a variety of needs that they seek to gratify by engaging with the courses, such as meeting new people and learning English. To meet these needs, we propose multiple design directions, including virtual social spaces outside any particular course, improved support for local groups of learners, and modularization to promote accessibility and organization of course content. Motivations thus provide a lens for understanding online learners and designing online courses to better support their needs.
René F. Kizilcec, Emily Schneider
ACM Trans. Comput. Hum. Interact.1
2014 Showing face in video instruction: effects on information retention, visual attention, and affect
abstract
The amount of online educational content is rapidly increasing, particularly in the form of video lectures. The goal is to design video instruction to facilitate an experience that maximizes learning and satisfaction. A widely used but understudied design element in video instruction is the overlay of a small video of the instructor over lecture slides. We conducted an experiment with eye-tracking and recall tests to investigate how adding the instructor's face to video instruction affects information retention, visual attention, and affect. Participants strongly preferred instruction with the face and perceived it as more educational. They spent about 41% of time looking at the face and switched between the face and slide every 3.7 seconds. Consistent with prior work, no significant difference in short- and medium-term recall ability was found. Including the face in video instruction is encouraged based on learners' positive affective response. More fine-grained analytics combining eye-tracking with detailed learning assessment could shed light on the mechanisms by which the face aids or hinders learning.
René F. Kizilcec, Kathryn Papadopoulos, Lalida Sritanyaratana
CHI1
2014 Anonymity in Social Media: Effects of Content Controversiality and Social Endorsement on Sharing Behavior
Kaiping Zhang, René F. Kizilcec
ICWSM2
2014 Reducing non-response bias with survey reweighting: applications for online learning researchers
abstract
In many online courses, information about learners is collected via surveys for accounting, instructional design, and research purposes. Aggregate information from such surveys is frequently reported in news articles and research papers, among other publications. While some authors acknowledge the potential bias due to non-response in course surveys, there are no investigations on the severity of the bias and methods for bias reduction in the online education context. A regression-based response-propensity model is described and applied to reweight a course survey, and discrepancies between adjusted and unadjusted outcome distributions are provided.
René F. Kizilcec
L@S1
2014 "Why did you enroll in this course?": developing a standardized survey question for reasons to enroll
abstract
Understanding motivations for enrolling in MOOCs is key for personalizing and scaling the online learning experience. We develop a standardized survey item for measuring learners' reasons to enroll, based on a corpus of open-ended responses from previous course surveys. Online coders were employed in the iterative development of response options. The item was designed to minimize response biases by adhering to best practices from survey design research.
Emily Schneider, René F. Kizilcec
L@S2
2013 Deconstructing disengagement: analyzing learner subpopulations in massive open online courses
abstract
As MOOCs grow in popularity, the relatively low completion rates of learners has been a central criticism. This focus on completion rates, however, reflects a monolithic view of disengagement that does not allow MOOC designers to target interventions or develop adaptive course features for particular subpopulations of learners. To address this, we present a simple, scalable, and informative classification method that identifies a small number of longitudinal engagement trajectories in MOOCs. Learners are classified based on their patterns of interaction with video lectures and assessments, the primary features of most MOOCs to date.
René F. Kizilcec, Chris Piech, Emily Schneider
LAK1