VLDB 2026 Research / reviewers in the wild / expert
Chris Piech
dblp:35/10987 · also Christopher Piech
· DBLP profile ↗
66ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0001-5140-0467ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 6 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 27 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 16 since 2021Systems, architecture and hardware · 15 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Large-Scale Observational Study on Obtaining Lightweight, Randomized Weekly Student Feedback: Associations with End-of-Term Course EvaluationsabstractConventional methods of obtaining student feedback on course experiences face a fundamental tradeoff between feedback frequency and quality: as feedback requests become more frequent, participation often declines, and responses become less thoughtful over time. To obtain both timely and thoughtful feedback from students, Kim and Piech [27] recently proposed a simple, lightweight course feedback mechanism: surveying each student a small number of times per term during randomly selected weeks. Termed High-Resolution Course Feedback (HRCF), this method has been shown to elicit feedback that instructors find helpful without placing excessive survey burden on students. Yunsung Kim, Candace Thille, Chris Piech |
L@S | 4 |
| 2026 | Personalized Exam Prep (PEP): Scaling No-Stakes, No-LLM Dialogue-Based Assessments in a Large CS CourseabstractAssessing what students know and what work they have done is increasingly challenging in the age of large language models (LLMs) like ChatGPT. At the same time, as class sizes have grown, the CS learning experience has become increasingly depersonalized. Dialogue-based assessments can simultaneously address both of these challenges; but although the benefits of dialogue-based assessments are well-documented, adopting these approaches is often too challenging in lower-division CS courses with hundreds of students. Thus, we sought to develop a dialogue-based intervention that augments existing course assessment structure, accounts for potential student LLM usage, scales to common CS class sizes, and benefits student learning via student-centered design principles. We report our experience implementing ''Personalized Exam Prep'' (PEP) in two offerings of a 300-student, second-year CS-core course at a North American research-intensive university. PEP consists of one-on-one, TA-guided, problem-solving dialogues, optimized to give students personalized feedback before an upcoming written exam. We describe the logistics of applying this intervention at scale, made possible by custom tooling. We then outline outcomes: positive student feedback, metacognitive impacts, how well PEP-measured performance predicts written exam scores, and the value of PEP as a course-diagnostic for instructors. Finally, we offer insights that may be valuable for large-course instructors updating their assessment structures in a post-LLM world. Kelly Cochran, Chris Piech |
SIGCSE (1) | 2 |
| 2026 | Aligning Small Language Models for Programming Feedback: Towards Scalable Coding Support in a Massive Global CourseabstractProviding timely and actionable feedback is essential for students learning to program. While large language models (LLMs) are increasingly used to automate this process, they remain costly to deploy and raise concerns around privacy and institutional control. Small language models (SLMs) offer a promising alternative: they can be run locally and integrated more flexibly into educational platforms. However, their out-of-the-box performance is often poor, requiring targeted training to be effective in classrooms. In this paper, we investigate whether a trained 3B-parameter SLM, guided by rubric-based prompting and a pipeline combining supervised and preference-based learning, can generate diagnostic feedback that approaches the quality of larger models. We deploy the model in a large-scale online programming course and compare its feedback to its base and fine-tuned variants, Llama-3.1-8B, and GPT-4.1, using human ratings from 53 teaching assistants and an automated LLM-as-a-judge analysis. Our results show that careful training narrows the feedback quality gap between an SLM and an LLM from over 80 to just 10 percentage points on key metrics. The trained SLM more rarely hallucinates errors, is often rated as helpful by educators, and only occasionally misses issues in student code. These findings suggest that small models can serve as practical and scalable targeted feedback solutions in large educational settings, while LLMs may remain necessary for more comprehensive diagnostic feedback. Charles Koutcheme, Juliette Woodrow, Chris Piech |
SIGCSE (1) | 3 |
| 2026 | Rooms of Their Own: Structured Small-Group Learning in a Realtime Browser-Based IDE
Jacob Roberts-Baca, Joshua Delgadillo, Chris Piech |
SIGCSE (1) | 3 |
| 2025 | AI Web Agents Can Effectively Guide Lesson Design and Predict Student Outcomes
Sierra Wang, John C. Mitchell, Chris Piech |
AIED (2) | 3 |
| 2025 | Improving Generative AI Student Feedback: Direct Preference Optimization with Teachers in the Loop
Juliette Woodrow, Chris Piech, Oluwasanmi Koyejo |
EDM | 2 |
| 2025 | Soft Grades: A Calibrated and Accurate Method for Course-Grade Estimation that Expresses UncertaintyabstractIn traditional educational settings, students are often summarized by a single number-a final course grade-that reflects their performance.While final grades are convenient for reporting or comparison, they oversimplify a student's true ability and do not express uncertainty.In this paper, we introduce a new item-response model for classroom settings that infers a distribution over student abilities and uses this to represent each student's final grade as a probability distribution.This approach captures the uncertainty that comes from variations in both student performance and grading processes.Practical applications of our approach include enabling teachers to better understand grading confidence, impute missing assignment scores, and make informed decisions when curving final grades.For students, the model offers probabilistic estimates of their final course grades based on current performance, supporting informed academic decisions such as opting for Pass/Fail grading.We evaluate our model using real-world datasets, showing that the Soft Grades model is well-calibrated and surpasses the state-of-the-art polytomous IRT model in accurately predicting future scores.Additionally, we share a web application and Python scripts to make our model available to teachers and students. Juliette Woodrow, Chris Piech |
LAK | 2 |
| 2025 | The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam PerformancesabstractLarge language models (LLMs) are quickly being adopted in a wide range of learning experiences, especially via ubiquitous and broadly accessible chat interfaces like ChatGPT. This type of interface is readily available to students and teachers around the world. Coding education is an interesting test case, both because LLMs have strong performance on coding tasks, and because LLM-powered support tools are rapidly becoming part of the workflow of professional software engineers. To help understand the impact of generic LLM use on coding education, we conducted a large-scale randomized control trial with 5,831 students from 146 countries in an online coding class in which we provided some students with access to a chat interface with GPT-4. Under some assumptions, we estimate positive benefits on exam performance for adopters, the students who used the tool, but over all students, the advertisement of GPT-4 led to a significant average decrease in exam participation. We observe similar decreases in other forms of course engagement. However, this decrease is modulated by the student's country of origin. Offering access to LLMs to students from low human development index countries increased their exam participation rate on average. Our results suggest there may be promising benefits to using LLMs in an introductory coding class, but also potential harms for engagement, which makes their longer term impact on student success unclear. Our work highlights the need for additional investigations to help understand the potential impact of future adoption and integration of LLMs into classrooms. Allen Nie, Yash Chandak, Miroslav Suzara, Ali Malik, Juliette Woodrow, Matt Peng, Mehran Sahami, Emma Brunskill, Chris Piech |
L@S | 9 |
| 2025 | The Next Educational Revolution: Grand Challenges for Learning @ Scale in the Generative AI EraabstractThe community working on learning at scale has made tremendous progress over the last decade, successfully achieving many of our previously stated grand challenges. As we enter the Generative AI era, what new ambitious milestones should we shoot for in order to make progress towards the joyful, high-quality education at scale for all learners? How can we get ahead of the curve of the disruption that could come to assessment and jobs? This talk will explore several potential objectives, including scaling human teaching, developing effective generative AI tools, reaching new heights in student understanding, and addressing a major persistent constraint: student motivation. Chris Piech |
L@S | 1 |
| 2025 | The Effects of Chatbot Placement, Personification, and Functionality on Student Outcomes in a Global CS1 Course
Sierra Wang, Thomas Jefferson, Chris Piech, John C. Mitchell |
L@S | 3 |
| 2025 | Fostering and Understanding Diverse Interpersonal Connections in a Massive Online CS1 CourseabstractForming social relationships is critical to student success and well-being, but is one of the first aspects to be neglected in the design of massive online courses. We present our experience deploying an in-course networking tool that enabled 1,600+ learners and teachers in a massive online CS1 course to form 2,000+ connections with other individuals. We discuss how social preferences and networking goals vary by demographics, economic factors, course goals, and course role. Contrary to usual online social behavior, users in our network sent more out-group requests than a random baseline by role (2.04x), gender (1.1x), and developing vs. developed country (1.07x). We highlight differences between developing vs. developed country users: developing country users send 2.5x requests and make, on average, 1.78x as many connections as those from developed countries. From a randomized control trial we find that random recommendations increase the volume of sent requests by 44.48% and promote cross-group requests across developing vs. developed countries (+28.9%), age (+15.1%), and gender (+8.6%). Ultimately we show that integrating socialization as a core feature of online CS1 classrooms can help support people from all backgrounds in achieving their diverse educational goals, which often extend well beyond improving coding proficiency. Miranda Li, Ali Malik, Chris Piech |
SIGCSE (1) | 3 |
| 2025 | Infinite Story
Chris Piech, Mehran Sahami, Yasmine Alonso, Katie Liu, Javokhir Arifov, Anjali Sreenivas, Dan Webber, Tina Zheng, Ngoc Nguyen, Iddah Mlauzi, Juliette Woodrow |
SIGCSE (2) | 1 |
| 2024 | Grading and Clustering Student Programs That Produce Probabilistic Output
Yunsung Kim, Jadon Geathers, Chris Piech |
EDM | 3 |
| 2024 | Faster Maximum Inner Product Search in High DimensionsabstractMaximum Inner Product Search (MIPS) is a ubiquitous task in machine learning applications. Given a query vector and $n$ other vectors in $d$ dimensions, the MIPS problem is to find the atom that has the highest inner product with the query vector. Existing MIPS algorithms scale at least as $O(\sqrt{d})$ with respect to $d$, which becomes computationally prohibitive in high-dimensional settings. In this work, we present BanditMIPS, a novel randomized algorithm that provably improves the state-of-the-art complexity from $O(\sqrt{d})$ to $O(1)$ with respect to $d$. We validate the scaling of BanditMIPS and demonstrate that BanditMIPS outperforms prior state-of-the-art MIPS algorithms in sample complexity, wall-clock time, and precision/speedup tradeoff across a variety of experimental settings. Furthermore, we propose a variant of our algorithm, named BanditMIPS-$\alpha$, which improves upon BanditMIPS by employing non-uniform sampling across coordinates. We also demonstrate the usefulness of BanditMIPS in problems for which MIPS is a subroutine, including Matching Pursuit and Fourier analysis. Finally, we demonstrate that BanditMIPS can be used in conjunction with preprocessing techniques to improve its complexity with respect to $n$. All of our experimental results are reproducible via a 1-line script at github.com/ThrunGroup/BanditMIPS. Mo Tiwari, Ryan Kang, Chris Piech, Sebastian Thrun, Ilan Shomorony, Martin J. Zhang |
ICML | 5 |
| 2024 | TeachNow: Enabling Teachers to Provide Spontaneous, Realtime 1: 1 Help in Massive Online CoursesabstractOne-on-one help from a teacher is highly impactful for students, yet extremely challenging to support in massive online courses (MOOCs). In this work, we present TeachNow: a novel system that lets volunteer teachers from anywhere in the world instantly provide 1:1 help sessions to students in MOOCs, without any scheduling or coordination overhead. TeachNow works by quickly finding an online student to help and putting them in a collaborative working session with the teacher. The spontaneous, on-demand nature of TeachNow gives teachers the flexibility to help whenever their schedule allows. Ali Malik, Juliette Woodrow, Chris Piech |
ITiCSE (1) | 4 |
| 2024 | Handwritten Code Recognition for Pen-and-Paper CS EducationabstractTeaching Computer Science (CS) by having students write programs by hand on paper has key pedagogical advantages: It allows focused learning and requires careful thinking compared to the use of Integrated Development Environments (IDEs) with intelligent support tools or "just trying things out". The familiar environment of pens and paper also lessens the cognitive load of students with no prior experience with computers, for whom the mere basic usage of computers can be intimidating. Finally, this teaching approach opens learning opportunities to students with limited access to computers. However, a key obstacle is the current lack of teaching methods and support software for working with and running handwritten programs. Optical character recognition (OCR) of handwritten code is challenging: Minor OCR errors, perhaps due to varied handwriting styles, easily make code not run, and recognizing indentation is crucial for languages like Python but is difficult to do due to inconsistent horizontal spacing in handwriting. Our approach integrates two innovative methods. The first combines OCR with an indentation recognition module and a language model designed for post-OCR error correction without introducing hallucinations. This method, to our knowledge, surpasses all existing systems in handwritten code recognition. It reduces error from 30% in the state of the art to 5% with minimal hallucination of logical fixes to student programs. The second method leverages a multimodal language model to recognize handwritten programs in an end-to-end fashion. We hope this contribution can stimulate further pedagogical research and contribute to the goal of making CS education universally accessible. We release a dataset of handwritten programs and code to support future research. Md Sazzad Islam, Moussa Doumbouya, Christopher D. Manning, Chris Piech |
L@S | 4 |
| 2024 | PyodideU: Unlocking Python Entirely in a Browser for CS1abstractIn this paper, we present an education-focused Python IDE and runtime library which can run entirely in desktop, laptop, tablet, and mobile device web browsers. Our solution provides features useful for an engaging CS1 course, and eliminates the need for a server-based runtime. We describe a new, open source, methodology for running interactive Python entirely in the browser by solving the "WebAssembly blocking problem," a core technical challenge to a web-based Python solution. Thomas Jefferson, Chris Gregg, Chris Piech |
SIGCSE (1) | 3 |
| 2024 | A Fast and Accurate Machine Learning Autograder for the Breakout AssignmentabstractIn this paper, we detail the successful deployment of a machine learning autograder that significantly decreases the grading labor required in the Breakout computer science assignment. This assignment - which tasks students with programming a game consisting of a controllable paddle and a ball that bounces off the paddle to break bricks - is popular for engaging students with introductory computer science concepts, but creates a large grading burden. Due to the game's interactive nature, grading defies traditional unit tests and instead typically requires 8+ minutes of manually playing each student's game to search for bugs. This amounts to 45+ hours of grading in a standard course offering and prevents further widespread adoption of the assignment. Our autograder alleviates this burden by playing each student's game with a reinforcement learning agent and providing videos of discovered bugs to instructors. In an A/B test with manual grading, we find that our human-in-the-loop AI autograder reduces grading time by 44%, while slightly improving grading accuracy by 6%, ultimately saving roughly 30 hours over our deployment in two offerings of the assignment. Our results further suggest the practicality of grading other interactive assignments (e.g., other games or building websites) via similar machine learning techniques. Live demo at https://ezliu.github.io/breakoutgrader. Evan Zheran Liu, David Yuan 0001, Elyse Cornwall, Juliette Woodrow, Kaylee Burns, Allen Nie, Emma Brunskill, Chris Piech, Chelsea Finn |
SIGCSE (1) | 9 |
| 2024 | Learners Teaching Novices: An Uplifting Alternative AssessmentabstractWe propose and carry-out a novel method of formative assessment called Assessment via Teaching (AVT), in which learners demonstrate their understanding of CS1 topics by tutoring more novice students. AVT has powerful benefits over traditional forms of assessment: it is centered around service to others and is highly rewarding for the learners who teach. Moreover, teaching greatly improves the learners' own understanding of the material and has a huge positive impact on novices, who receive free 1:1 tutoring. Lastly, this form of assessment is naturally difficult to cheat---a critical property for assessments in the era of large-language models. We use AVT in a randomised control trial with learners in a CS1 course at an R1 university. The learners provide tutoring sessions to more novice students taking a lagged online version of the same course. We show that learners who do an AVT session before the course exam performed 20 to 30 percentage points better than the class average on several questions. Moreover, compared to students who did a practice exam, the AVT learners enjoyed their experience more and were twice as likely to study for their teaching session. We believe AVT is a scalable and uplifting method for formative assessment that could one day replace traditional exams. Ali Malik, Juliette Woodrow, Chris Piech |
SIGCSE (1) | 3 |
| 2024 | Math IDE: A Platform for Creating with MathabstractTo inspire student engagement in middle school math, we explore the possibility of using generative AI to enhance the creativity of math learning. We present the Math IDE, a math education environment in which students learn about math concepts by building artifacts. We aimed to create a platform in which students can engage with mathematical concepts, create an artifact that embodies the math that they are learning about, and practice their high-level specification skills. In the current iteration of the Math IDE, students can create custom web pages by describing and demonstrating understanding of the math that is involved in the web page. In this short overview, we describe our process and discuss several open questions regarding the design and application of this novel method of math education. Sierra Wang, John C. Mitchell, Nick Haber, Chris Piech |
SIGCSE (2) | 4 |
| 2024 | A Large Scale RCT on Effective Error Messages in CS1abstractIn this paper, we evaluate the most effective error message types through a large-scale randomized controlled trial conducted in an open-access, online introductory computer science course with 8,762 students from 146 countries. We assess existing error message enhancement strategies, as well as two novel approaches of our own: (1) generating error messages using OpenAI's GPT in real time and (2) constructing error messages that incorporate the course discussion forum. By examining students' direct responses to error messages, and their behavior throughout the course, we quantitatively evaluate the immediate and longer term efficacy of different error message types. We find that students using GPT generated error messages repeat an error 23.1% less often in the subsequent attempt, and resolve an error in 34.8% fewer additional attempts, compared to students using standard error messages. We also perform an analysis across various demographics to understand any disparities in the impact of different error message types. Our results find no significant difference in the effectiveness of GPT generated error messages for students from varying socioeconomic and demographic backgrounds. Our findings underscore GPT generated error messages as the most helpful error message type, especially as a universally effective intervention across demographics. Sierra Wang, John C. Mitchell, Chris Piech |
SIGCSE (1) | 3 |
| 2024 | AI Teaches the Art of Elegant Coding: Timely, Fair, and Helpful Style Feedback in a Global CourseabstractTeaching students how to write code that is elegant, reusable, and comprehensible is a fundamental part of CS1 education. However, providing this "style feedback" in a timely manner has proven difficult to scale. In this paper, we present our experience deploying a novel, real-time style feedback tool in Code in Place, a large-scale online CS1 course. Our tool is based on the latest breakthroughs in large-language models (LLMs) and was carefully designed to be safe and helpful for students. We used our Real-Time Style Feedback tool (RTSF) in a class with over 8,000 diverse students from across the globe and ran a randomized control trial to understand its benefits. We show that students who received style feedback in real-time were five times more likely to view and engage with their feedback compared to students who received delayed feedback. Moreover, those who viewed feedback were more likely to make significant style-related edits to their code, with over 79% of these edits directly incorporating their feedback. We also discuss the practicality and dangers of LLM-based tools for feedback, investigating the quality of the feedback generated, LLM limitations, and techniques for consistency, standardization, and safeguarding against demographic bias, all of which are crucial for a tool utilized by students. Juliette Woodrow, Ali Malik, Chris Piech |
SIGCSE (1) | 3 |
| 2023 | Variational Temporal IRT: Fast, Accurate, and Explainable Inference of Dynamic Learner Proficiency
Yunsung Kim, Sreechan Sankaranarayanan, Chris Piech, Candace Thille |
EDM | 3 |
| 2023 | The Student Zipf Theory: Inferring Latent Structures in Open-Ended Student Work To Help EducatorsabstractAre there structures underlying student work that are universal across every open-ended task? We demonstrate that, across many subjects and assignment types, the probability distribution underlying student-generated open-ended work is close to Zipf’s Law. Inferring this latent structure for classroom assignments can help learning analytics researchers, instruction designers, and educators understand the landscape of various student approaches, assess the complexity of assignments, and prioritise pedagogical attention. However, typical classrooms are way too small to witness even the contour of the Zipfian pattern, and it is generally impossible to perform inference for Zipf’s law from such small number of samples. We formalise this difficult task as the Zipf Inference Challenge: (1) Infer the ordering of student-generated works by their underlying probabilities, and (2) Estimate the shape parameter of the underlying distribution in a typical-sized classroom. Our key insight in addressing this challenge is to leverage the densities of the student response landscapes represented by semantic similarity. We show that our “Semantic Density Estimation” method is able to do a much better job at inferring the latent Zipf shape and the probability-ordering of student responses for real world education datasets. Yunsung Kim, Chris Piech |
LAK | 2 |
| 2023 | High-Resolution Course Feedback: Timely Feedback Mechanism for InstructorsabstractWe study the problem of minimizing the delay between when an issue comes up in a course and when instructors get feedback about it. The widespread practice of obtaining midterm and end-of-term feedback from students is suboptimal in this regard, especially for large courses: it over-samples at a specific point in the course and can be biased by factors irrelevant to the teaching process. As a solution, we release High Resolution Course Feedback (HRCF), an open-source student feedback mechanism that builds on a surprisingly simple idea: survey each student on random weeks exactly twice per term. Despite the simplicity of its core idea, when deployed to 31 courses totaling a cumulative 6,835 students, HRCF was able to detect meaningful mood changes in courses and significantly improve timely feedback without asking for extra work from students compared to the common practice. An interview with the instructors revealed that HRCF provided constructive and useful feedback about their courses early enough to be acted upon, which would have otherwise been unobtainable through other survey methods. We also explore the possibility of using Large Language Models to flexibly and intuitively organize large volumes of student feedback at scale and discuss how HRCF can be further improved. Yunsung Kim, Chris Piech |
L@S | 2 |
| 2023 | GPTeach: Interactive TA Training with GPT-based StudentsabstractInteractive and realistic teacher training is hard to scale. This is a key issue for learning at scale, as inadequate preparation can negatively impact both students and teachers. What if we could make the teacher training experience more engaging and, as a downstream effect, reduce the potential for harm that teachers-in-training could inflict on students? We present GPTeach, an interactive chat-based teacher training tool that allows novice teachers to practice with simulated students. We performed two studies to evaluate GPTeach: one think-aloud study and one A/B test between our tool and a baseline. Participants took the role of a teaching assistant conducting office hours with two GPT-simulated students. We found that our tool provides the opportunity for teachers to get valuable teaching practice without the pressures of affecting real students, allowing them to iterate their responses both during and across sessions. Additionally, participants enjoyed flexibility in tailoring their responses according to the varied personas, needs, and learning goals. In this paper, we provide quantitative results and qualitative observations to inform future work in this area. We conclude with a discussion of actionable design ideas for such systems, as well as other ways to use this tool for evaluating teachers and students. GPTeach has recently been deployed into the teacher training component of an online course with over 800 novice teachers. Julia M. Markel, Steven G. Opferman, James A. Landay, Chris Piech |
L@S | 4 |
| 2023 | MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment TasksabstractHuman commonsense understanding of the physical and social world is organized around intuitive theories. These theories support making causal and moral judgments. When something bad happens, we naturally ask: who did what, and why? A rich literature in cognitive science has studied people's causal and moral intuitions. This work has revealed a number of factors that systematically influence people's judgments, such as the violation of norms and whether the harm is avoidable or inevitable. We collected a dataset of stories from 24 cognitive science papers and developed a system to annotate each story with the factors they investigated. Using this dataset, we test whether large language models (LLMs) make causal and moral judgments about text-based scenarios that align with those of human participants. On the aggregate level, alignment has improved with more recent LLMs. However, using statistical analyses, we find that LLMs weigh the different factors quite differently from human participants. These results show how curated, challenge datasets combined with insights from cognitive science can help us go beyond comparisons based merely on aggregate metrics: we uncover LLMs implicit tendencies and show to what extent these align with human intuitions. Allen Nie, Atharva Amdekar, Chris Piech, Tatsunori B. Hashimoto, Tobias Gerstenberg |
NeurIPS | 4 |
| 2023 | Detecting the Reasons for Program Decomposition in CS1 and Evaluating Their ImpactabstractDecomposition is considered one of the four cornerstones of computational thinking, which is essential to software development [36]. It requires the ability to assess a problem at a high level, develop a strategy to combat it, and then design a solution. Our study focuses on the metacognitive aspect of decomposition. We try to understand the learner's thought process and, specifically, what makes the novice programmer decide to break down a function. Charis Charitsis, Chris Piech, John C. Mitchell |
SIGCSE (1) | 2 |
| 2022 | The AI Teacher Test: Measuring the Pedagogical Ability of Blender and GPT-3 in Educational Dialogues
Anaïs Tack, Chris Piech |
EDM | 2 |
| 2022 | Function Names: Quantifying the Relationship Between Identifiers and Their Functionality to Improve ThemabstractWhen students first learn to program, they often focus on functionality: does a program work? In an era where software volume and complexity increase exponentially, it is equally important that they learn to write programs with style so that they are readable and extendable. Writing quality code starts with the building blocks for any program, its functions. A carefully chosen name is vital for program maintainability and manageability. The identifier is the most portable and concise way to summarize what the function does. What makes for the right choice? And can we automatically assess the quality of function names? Using natural language processing, we were able to create a probabilistic model to evaluate their clarity. Using functionality encodings, we attempt to learn the relationship between functions in different programs to improve their names. We analyzed a total of 5,400 programs tackling five novice programming tasks submitted by over 1,000 students in CS1. We developed a software system to automate labor-intensive tasks, detect poor function names and recommend replacements. Our findings suggest that less than 2.5% of name substitutions have an adverse outcome, and in most cases, more than 50% result in an improvement. Charis Charitsis, Chris Piech, John C. Mitchell |
L@S | 2 |
| 2022 | Using NLP to Quantify Program Decomposition in CS1abstractDecomposition is a problem-solving technique that is essential to software development. Nonetheless, it is perceived as the most challenging programming skill for learners to master. Researchers have studied decomposition in introductory programming courses through guided experiments, case studies, and surveys. We believe that the rapid advancements in scientific fields such as machine learning and natural language processing (NLP) opened up opportunities for more scalable approaches. Charis Charitsis, Chris Piech, John C. Mitchell |
L@S | 2 |
| 2022 | Giving Feedback on Interactive Student Programs with Meta-ExplorationabstractDeveloping interactive software, such as websites or games, is a particularly engaging way to learn computer science. However, teaching and giving feedback on such software is time-consuming — standard approaches require instructors to manually grade student-implemented interactive programs. As a result, online platforms that serve millions, like Code.org, are unable to provide any feedback on assignments for implementing interactive programs, which critically hinders students’ ability to learn. One approach toward automatic grading is to learn an agent that interacts with a student’s program and explores states indicative of errors via reinforcement learning. However, existing work on this approach only provides binary feedback of whether a program is correct or not, while students require finer-grained feedback on the specific errors in their programs to understand their mistakes. In this work, we show that exploring to discover errors can be cast as a meta-exploration problem. This enables us to construct a principled objective for discovering errors and an algorithm for optimizing this objective, which provides fine-grained feedback. We evaluate our approach on a set of over 700K real anonymized student programs from a Code.org interactive assignment. Our approach provides feedback with 94.3% accuracy, improving over existing approaches by 17.7% and coming within 1.5% of human-level accuracy. Project web page: https://ezliu.github.io/dreamgrader. Evan Zheran Liu, Moritz Stephan, Allen Nie, Chris Piech, Emma Brunskill, Chelsea Finn |
NeurIPS | 4 |
| 2022 | MABSplit: Faster Forest Training Using Multi-Armed BanditsabstractRandom forests are some of the most widely used machine learning models today, especially in domains that necessitate interpretability. We present an algorithm that accelerates the training of random forests and other popular tree-based learning methods. At the core of our algorithm is a novel node-splitting subroutine, dubbed MABSplit, used to efficiently find split points when constructing decision trees. Our algorithm borrows techniques from the multi-armed bandit literature to judiciously determine how to allocate samples and computational power across candidate split points. We provide theoretical guarantees that MABSplit improves the sample complexity of each node split from linear to logarithmic in the number of data points. In some settings, MABSplit leads to 100x faster training (an 99% reduction in training time) without any decrease in generalization performance. We demonstrate similar speedups when MABSplit is used across a variety of forest-based variants, such as Extremely Random Forests and Random Patches. We also show our algorithm can be used in both classification and regression tasks. Finally, we show that MABSplit outperforms existing methods in generalization performance and feature importance calculations under a fixed computational budget. All of our experimental results are reproducible via a one-line script at https://github.com/ThrunGroup/FastForest. Mo Tiwari, Ryan Kang, Chris Piech, Ilan Shomorony, Sebastian Thrun, Martin J. Zhang |
NeurIPS | 4 |
| 2022 | Feedback on Program Development Process for CS1 StudentsabstractIn introductory CS programming courses, student learning is often assessed on the basis of the submitted code. However, the final artifact fails to capture the problem-solving journey. Active learning occurs when a student stumbles upon conceptual unclarities, design dilemmas, algorithmic challenges. How can we shine light on hidden programming aspects to help teachers provide insightful feedback to learners? We developed a tool to analyze programs snapshots from the entire development process, visualize their evolution over time, filter syntax errors, and detect complex code. The tool is equipped with an editor to explore different ideas and execute the program on the fly, especially in one-on-one feedback sessions with the learner. We are eager to share our work with researchers and educators in other institutions and look forward to their feedback and ideas for improvement. Charis Charitsis, Chris Piech, John C. Mitchell |
SIGCSE (2) | 2 |
| 2021 | Using Radio Archives for Low-Resource Speech Recognition: Towards an Intelligent Virtual Assistant for Illiterate UsersabstractFor many of the 700 million illiterate people around the world, speech recognition technology could provide a bridge to valuable information and services. Yet, those most in need of this technology are often the most underserved by it. In many countries, illiterate people tend to speak only low-resource languages, for which the datasets necessary for speech technology development are scarce. In this paper, we investigate the effectiveness of unsupervised speech representation learning on noisy radio broadcasting archives, which are abundant even in low-resource languages. We make three core contributions. First, we release two datasets to the research community. The first, West African Radio Corpus, contains 142 hours of audio in more than 10 languages with a labeled validation subset. The second, West African Virtual Assistant Speech Recognition Corpus, consists of 10K labeled audio clips in four languages. Next, we share West African wav2vec, a speech encoder trained on the noisy radio corpus, and compare it with the baseline Facebook speech encoder trained on six times more data of higher quality. We show that West African wav2vec performs similarly to the baseline on a multilingual speech recognition task, and significantly outperforms the baseline on a West African language identification task. Finally, we share the first-ever speech recognition models for Maninka, Pular and Susu, languages spoken by a combined 10 million people in over seven countries, including six where the majority of the adult population is illiterate. Our contributions offer a path forward for ethical AI research to serve the needs of those most disadvantaged by the digital divide. Moussa Doumbouya, Lisa Einstein, Chris Piech |
AAAI | 3 |
| 2021 | SimGrade: Using Code Similarity Measures for More Accurate Human Grading
Sonja Johnson-Yu, Nicholas Bowman, Mehran Sahami, Chris Piech |
EDM | 4 |
| 2021 | Generative Grading: Near Human-level Accuracy for Automated Feedback on Richly Structured Problems
Ali Malik, Mike Wu, Vrinda Vasavada, Jinpeng Song, Madison Coots, Noah D. Goodman, Chris Piech |
EDM | 8 |
| 2021 | Assessing Function Names and Quantifying the Relationship Between Identifiers and Their Functionality to Improve ThemabstractWhen students first learn to program, they often focus on functionality: does a program work? In an era where software volume and complexity increase exponentially, it is equally important that they learn to write code with style. Quality code starts with the building blocks for any program, its functions. A carefully chosen name is vital for program maintainability and manageability. The identifier is the most portable and concise way to summarize what the function does. What makes for the right choice? And can we automatically assess the quality of function names? Using natural language processing, we were able to create a probabilistic model to evaluate their clarity. Using functionality encodings, we attempt to learn the relationship between functions in different programs to improve their names. We analyzed a total of 3,900 programs tackling three novice programming tasks submitted by 1,300 students in CS1. Charis Charitsis, Chris Piech, John C. Mitchell |
L@S | 2 |
| 2021 | Play to Grade: Testing Coding Games as Classifying Markov Decision ProcessabstractContemporary coding education often presents students with the task of developing programs that have user interaction and complex dynamic systems, such as mouse based games. While pedagogically compelling, there are no contemporary autonomous methods for providing feedback. Notably, interactive programs are impossible to grade by traditional unit tests. In this paper we formalize the challenge of providing feedback to interactive programs as a task of classifying Markov Decision Processes (MDPs). Each student's program fully specifies an MDP where the agent needs to operate and decide, under reasonable generalization, if the dynamics and reward model of the input MDP should be categorized as correct or broken. We demonstrate that by designing a cooperative objective between an agent and an autoregressive model, we can use the agent to sample differential trajectories from the input MDP that allows a classifier to determine membership: Play to Grade. Our method enables an automatic feedback system for interactive code assignments. We release a dataset of 711,274 anonymized student submissions to a single assignment with hand-coded bug labels to support future research. Allen Nie, Emma Brunskill, Chris Piech |
NeurIPS | 3 |
| 2021 | PearProgram: A More Fruitful Approach to Pair ProgrammingabstractIn this paper we present PearProgram, a hybrid learning and research tool that helps introductory Computer Science (CS) students learn how to pair program, including in remote learning environments. Grounded in theory from the Learning Sciences, the tool -- a collaborative, online IDE -- has two primary goals: 1) to help introductory CS students achieve pair programming success; and 2) to research what factors contribute to pairs that have beneficial outcomes. We present our learnings from the use of PearProgram in three remote introductory CS courses: a CS1 course, and two large international courses, including one for high school students. Teacher and student users responded positively to PearProgram, and use of the tool was associated with beneficial learning outcomes in these online learning environments. Our research opens many future research directions for (remote) pair programming, and indicates practices that may prove useful for CS educators at all levels. Maxwell Bigman, Ethan Roy, Miroslav Suzara, Chris Piech |
SIGCSE | 6 |
| 2021 | Code in Place: Online Section Leading for Scalable Human-Centered LearningabstractCould it be the case that the number of people who want to teach computer science, and have the potential, is roughly proportional to the number of people who want to learn? During the time of COVID-19 we offered a free CS1 class to people around the world. Well-aware of the high drop-out rates reported in many massive open-access online courses (MOOCs), we augmented our course with a scalable, human-centered solution: section leading. Section leaders teach small, weekly interactive learning sessions. We hypothesize that the personalized attention adds a sense of responsibility for both student and teacher which drives learning. We recruited over 900 volunteer section leaders and more than 10,000 students in the class. To our knowledge this is the largest group of section leaders in a single CS1 course offering and the most small group interactions. The completion rate in our class was more than 10 times that usually reported for similar MOOCs. Additionally, 99% of the volunteer section leaders taught through the entire span of the course, showing the potential for large scale volunteer-driven education, and the benefit that teachers themselves derive. We also discovered the potential for replication of this model, as 34% of students in a representative-sample survey indicated they would serve as section leaders for a future offering of the course. This level of participation would be more than sufficient to field additional offerings of the course sustainably. We believe this is an intriguing case study of a model for significantly scaling human-centric CS education for all. Chris Piech, Ali Malik, Kylie Jue, Mehran Sahami |
SIGCSE | 1 |
| 2020 | The Stanford Acuity Test: A Precise Vision Test Using Bayesian Techniques and a Discovery in Human Visual ResponseabstractChart-based visual acuity measurements are used by billions of people to diagnose and guide treatment of vision impairment. However, the ubiquitous eye exam has no mechanism for reasoning about uncertainty and as such, suffers from a well-documented reproducibility problem. In this paper we make two core contributions. First, we uncover a new parametric probabilistic model of visual acuity response based on detailed measurements of patients with eye disease. Then, we present an adaptive, digital eye exam using modern artificial intelligence techniques which substantially reduces acuity exam error over existing approaches, while also introducing the novel ability to model its own uncertainty and incorporate prior beliefs. Using standard evaluation metrics, we estimate a 74% reduction in prediction error compared to the ubiquitous chart-based eye exam and up to 67% reduction compared to the previous best digital exam. For patients with eye disease, the novel ability to finely measure acuity from home could be a crucial part in early diagnosis. We provide a web implementation of our algorithm for anyone in the world to use. The insights in this paper also provide interesting implications for the field of psychometric Item Response Theory. Chris Piech, Ali Malik, Laura M. Scott, Robert T. Chang, Charles Lin |
AAAI | 1 |
| 2020 | Measuring Ability-to-Learn Using Parametric Learning Gain Functions
Chris Piech, Engin Bumbacher, Richard Lee Davis |
EDM | 1 |
| 2020 | Variational Item Response Theory: Fast, Accurate, and Expressive
Mike Wu, Richard Lee Davis, Benjamin W. Domingue, Chris Piech, Noah D. Goodman |
EDM | 4 |
| 2020 | Using Google Search Trends to Estimate Global Patterns in LearningabstractThe use of the Internet for learning provides a unique and growing opportunity to revisit the task of quantifying how much people have learned about a given subject in different regions around the world. Google alone receives over 5 billion searches a day and its publicly available data provides insight into learning process that is otherwise unobservable on a global scale. In this paper we, introduce the Computer Science Literacy-Proxy Index via Search (CSLI-s), a measure that utilizes online search data to make an educated guess around trends in computer science education. This measure uses a statistical signal processing technique to compose search volumes from a spectrum of topics into a coherent score. We intentionally explore and mitigate the biases of search data and, in the process, develop CSLI-s scores that correlate with traditional, more expensive metrics of learning. We then use search-trend data to measure patterns in subject literacy across countries and over time. To the best of our knowledge, this is the first measure of learning via Internet search-trends. The Internet is becoming a standard tool for learners and, as such, we anticipate search-trend data will have growing relevance to the learning science community. Serhat Arslan, Mo Tiwari, Chris Piech |
L@S | 3 |
| 2020 | Human Languages in Source Code: Auto-Translation for Localized InstructionabstractComputer science education has promised open access around the world, but access is largely determined by what human language you speak. As younger students learn computer science it is less appropriate to assume that they should learn English beforehand. To that end, we present CodeInternational, the first tool to translate code between human languages. To develop a theory of non-English code, and inform our translation decisions, we conduct a study of public code repositories on GitHub. The study is to the best of our knowledge the first on human-language in code and covers 2.9 million Java repositories. To demonstrate CodeInternational's educational utility, we build an interactive version of the popular English-language Karel reader and translate it into 100 spoken languages. Our translations have already been used in classrooms around the world, and represent a first step in an important open CS-education problem. Chris Piech, Sami Abu-El-Haija |
L@S | 1 |
| 2020 | Co-Teaching Computer Science Across Borders: Human-Centric Learning at ScaleabstractProgramming is fast becoming a required skill set for students in every country. We present CS Bridge, a model for cross-border co-teaching of CS1, along with a corresponding open-source course-in-a-box curriculum made for easy localization. In the CS Bridge model, instructors and student-teachers from different countries come together to teach a short, stand-alone CS1 course to hundreds of local high school students. The corresponding open-source curriculum has been specifically designed to be easily adapted to a wide variety of local teaching practices, languages, and cultures. Chris Piech, Lisa Yan, Lisa Einstein, Ana Saavedra, Baris Bozkurt, Eliska Sestáková, Ondrej Guth, Nick McKeown |
L@S | 1 |
| 2020 | BanditPAM: Almost Linear Time k-Medoids Clustering via Multi-Armed BanditsabstractClustering is a ubiquitous task in data science. Compared to the commonly used k-means clustering, k-medoids clustering requires the cluster centers to be actual data points and supports arbitrary distance metrics, which permits greater interpretability and the clustering of structured objects. Current state-of-the-art k-medoids clustering algorithms, such as Partitioning Around Medoids (PAM), are iterative and are quadratic in the dataset size n for each iteration, being prohibitively expensive for large datasets. We propose BanditPAM, a randomized algorithm inspired by techniques from multi-armed bandits, that reduces the complexity of each PAM iteration from O(n^2) to O(nlogn) and returns the same results with high probability, under assumptions on the data that often hold in practice. As such, BanditPAM matches state-of-the-art clustering loss while reaching solutions much faster. We empirically validate our results on several large real-world datasets, including a coding exercise submissions dataset from Code.org, the 10x Genomics 68k PBMC single-cell RNA sequencing dataset, and the MNIST handwritten digits dataset. In these experiments, we observe that BanditPAM returns the same results as state-of-the-art PAM-like algorithms up to 4x faster while performing up to 200x fewer distance computations. The improvements demonstrated by BanditPAM enable k-medoids clustering on a wide range of applications, including identifying cell types in large-scale single-cell data and providing scalable feedback for students learning computer science online. We also release highly optimized Python and C++ implementations of our algorithm. Mo Tiwari, Martin J. Zhang, James Mayclin, Sebastian Thrun, Chris Piech, Ilan Shomorony |
NeurIPS | 5 |
| 2019 | Zero Shot Learning for Code Education: Rubric Sampling with Deep Learning InferenceabstractIn modern computer science education, massive open online courses (MOOCs) log thousands of hours of data about how students solve coding challenges. Being so rich in data, these platforms have garnered the interest of the machine learning community, with many new algorithms attempting to autonomously provide feedback to help future students learn. But what about those first hundred thousand students? In most educational contexts (i.e. classrooms), assignments do not have enough historical data for supervised learning. In this paper, we introduce a human-in-the-loop “rubric sampling” approach to tackle the “zero shot” feedback challenge. We are able to provide autonomous feedback for the first students working on an introductory programming assignment with accuracy that substantially outperforms data-hungry algorithms and approaches human level fidelity. Rubric sampling requires minimal teacher effort, can associate feedback with specific parts of a student’s solution and can articulate a student’s misconceptions in the language of the instructor. Deep learning inference enables rubric sampling to further improve as more assignment specific student data is acquired. We demonstrate our results on a novel dataset from Code.org, the world’s largest programming education platform. Mike Wu, Milan Mossé, Noah D. Goodman, Chris Piech |
AAAI | 4 |
| 2019 | Grades are not Normal: Improving Exam Score Models Using the Logit-Normal Distribution
Noah Arthurs, Ben Stenhaug, Sergey Karayev, Chris Piech |
EDM | 4 |
| 2019 | Using Latent Variable Models to Observe Academic Pathways
Nate Gruver, Ali Malik, Brahm Capoor, Chris Piech, Mitchell L. Stevens, Andreas Paepcke |
EDM | 4 |
| 2019 | Pensieve: Feedback on Coding Process for NovicesabstractIn large undergraduate computer science classrooms, student learning on assignments is often gauged only by the work on their final solution, not by their programming process. As a consequence, teachers are unable to give detailed feedback on how students implement programming methodology, and novice students often lack a metacognitive understanding of how they learn. We introduce Pensieve as a drag-and-drop, open-source tool that organizes snapshots of student code as they progress through an assignment. The tool is designed to encourage sit-down conversations between student and teacher about the programming process. The easy visualization of code evolution over time facilitates the discussion of intermediate work and progress towards learning goals, both of which would otherwise be unapparent from a single final submission. This paper discusses the pedagogical foundations and technical details of Pensieve and describes results from a particular 207-student classroom deployment, suggesting that the tool has meaningful impacts on education for both the student and the teacher. Lisa Yan, Annie Hu, Chris Piech |
SIGCSE | 3 |
| 2019 | The PyramidSnapshot Challenge: Understanding Student Process from Visual Output of ProgramsabstractIn the ideal CS1 classroom, we should understand programming process---how student code evolves over time. However, for graphics-based programming assignments, the task of understanding and grading final solutions, let alone thousands of intermediate steps, is incredibly labor-intensive. In this work, we present a challenge, a dataset, and a promising first solution to autonomously use image output to identify functional, intermediate stages of a student solution. By using computer vision techniques to associate visual output of intermediate student code with functional progress, we supplement a lot of the teacher labor associated with understanding graphics-based, open-ended assignments. We hope our publication of the dataset used in our case study sparks discussion in the community on how to analyze programs with visual program output. Lisa Yan, Nick McKeown, Chris Piech |
SIGCSE | 3 |
| 2018 | BlueBook: A Computerized Replacement for Paper Tests in Computer ScienceabstractThis paper presents BlueBook, a lightweight, cross-platform, computer-based, open source examination environment that overcomes traditional hurdles with computerized testing for computer science courses. As opposed to paper exam testing, BlueBook allows students to type coding problems on their laptops in an environment similar to their normal programming routine (e.g., with syntax highlighting), but purposefully does not provide them the ability to compile and/or run their code. We seamlessly transitioned from paper exams to BlueBook and found that students appreciated the ability to type their responses. Additionally, we are just beginning to harness the benefits to grading by having student answers in digital form. In the paper, we discuss the pedagogical benefits and trade-offs to using a computerized exam format, and we argue that both the students and the graders benefit from it. Chris Piech, Chris Gregg |
SIGCSE | 1 |
| 2018 | TMOSS: Using Intermediate Assignment Work to Understand Excessive Collaboration in Large ClassesabstractAs computer science classes grow, instructor workload also increases: teachers must simultaneously teach material, provide assignment feedback, and monitor student progress. At scale, it is hard to know which students need extra help, and as a result some students can resort to excessive collaboration--using online resources or peer code--to complete their work. In this paper, we present TMOSS, a tool that analyzes the intermediate steps a student takes to complete a programming assignment. We find that for three separate course offerings, TMOSS is almost twice as effective as traditional software similarity detectors in identifying the number of students who exhibit excessive collaboration. We also find that such students spend significantly less time on their assignment, use fewer class tutoring resources, and perform worse on exams than their peers. Finally, we provide a theory of the parametric distribution of typical student assignment similarity, which allows for probabilistic interpretation. Lisa Yan, Nick McKeown, Mehran Sahami, Chris Piech |
SIGCSE | 4 |
| 2017 | Learning to Represent Student Knowledge on Programming Exercises Using Deep Learning
Lisa Wang, Angela Sy, Larry Liu, Chris Piech |
EDM | 4 |
| 2017 | Deep Knowledge Tracing On Programming ExercisesabstractModeling a student's knowledge state while she is solving exercises is a crucial stepping stone towards providing better personalized learning experiences at scale. This task, also referred to as "knowledge tracing", has been explored extensively on exercises where student submissions fall into a finite discrete solution space, e.g. a multiple-choice answer. However, we believe that rich information about a student's learning is captured within their responses to open-ended problems with unbounded solution spaces, such as programming exercises. In addition, sequential snapshots of a student's progress while she is solving a single exercise can provide valuable insights into her learning behavior. In this setting, creating representations for a student's knowledge state is a challenging task, but with recent advances in machine learning, there are more promising techniques to learn representations for complex entities. In our work, we feed the embedded program submissions into a recurrent neural network and train it on the task of predicting the student's success on the subsequent programming exercise. By training on this task, the model learns nuanced representations of a student's knowledge, and reliably predicts future student performance. Lisa Wang, Angela Sy, Larry Liu, Chris Piech |
L@S | 4 |
| 2017 | Sharing and Using Programming Log Data (Abstract Only)abstractAs more programming environments add logging features and programming data becomes more accessible, it is important to have a conversation about how we share and use this data. Uses of programming log data range from big-picture analyses to dashboards for instant teacher feedback, to intelligent, data-driven learning environments. The goal of this BOF is to talk about what data is important to collect, where it can be gathered and shared, what general data formats make sense, how to handle privacy and anonymization, and what ultimately we want to see the data used for. The BOF welcomes both producers of programming log data and current or potential consumers, interested in how it could be applied in their classrooms or research. One hopeful outcome of this BOF is a commitment to documenting and sharing existing programming data in an accessible location and format. Thomas W. Price, Neil Brown 0001, Chris Piech, Kelly Rivers |
SIGCSE | 3 |
| 2016 | As CS Enrollments Grow, Are We Attracting Weaker Students?abstractIn recent years, enrollments in undergraduate computer science programs have seen tremendous growth nationally. Often accompanying such growth is a concern from faculty that the additional students choosing to pursue computing may not have the same aptitude for the subject as was seen in prior student populations. Thus such students may exhibit weaker performance in computing courses. To help address this question, we present a statistical analysis using mixture modeling of students' performance in an introductory programming class at Stanford University over an eight year period, during which enrollments in the course more than doubled. Importantly, in this setting many variables that would normally confound such a study are directly controlled for. We find that the distribution of student performance during this period, as reflected in their programming assignment scores, remains remarkably stable despite the large growth in enrollment. We then explain how the notion of having "more weak students" and the fact that the distribution of student ability is unchanged can readily co-exist and lead to misperceptions about the quality of incoming students during an enrollment boom. Mehran Sahami, Chris Piech |
SIGCSE | 2 |
| 2015 | Learning Program Embeddings to Propagate Feedback on Student CodeabstractProviding feedback, both assessing final work and giving hints to stuck students, is difficult for open-ended assignments in massive online classes which can range from thousands to millions of students. We introduce a neural network method to encode programs as a linear mapping from an embedded precondition space to an embedded postcondition space and propose an algorithm for feedback at scale using these linear maps as features. We apply our algorithm to assessments from the Code.org Hour of Code and Stanford University’s CS1 course, where we propagate human comments on student assignments to orders of magnitude more submissions. Chris Piech, Jonathan Huang, Andy Nguyen, Mike Phulsuksombati, Mehran Sahami, Leonidas J. Guibas |
ICML | 1 |
| 2015 | Autonomously Generating Hints by Inferring Problem Solving PoliciesabstractExploring the whole sequence of steps a student takes to produce work, and the patterns that emerge from thousands of such sequences is fertile ground for a richer understanding of learning. In this paper we autonomously generate hints for the Code.org `Hour of Code,' (which is to the best of our knowledge the largest online course to date) using historical student data. We first develop a family of algorithms that can predict the way an expert teacher would encourage a student to make forward progress. Such predictions can form the basis for effective hint generation systems. The algorithms are more accurate than current state-of-the-art methods at recreating expert suggestions, are easy to implement and scale well. We then show that the same framework which motivated the hint generating algorithms suggests a sequence-based statistic that can be measured for each learner. We discover that this statistic is highly predictive of a student's future success. Chris Piech, Mehran Sahami, Jonathan Huang, Leonidas J. Guibas |
L@S | 1 |
| 2015 | Deep Knowledge TracingabstractKnowledge tracing, where a machine models the knowledge of a student as they interact with coursework, is an established and significantly unsolved problem in computer supported education.In this paper we explore the benefit of using recurrent neural networks to model student learning.This family of models have important advantages over current state of the art methods in that they do not require the explicit encoding of human domain knowledge,and have a far more flexible functional form which can capture substantially more complex student interactions.We show that these neural networks outperform the current state of the art in prediction on real student data,while allowing straightforward interpretation and discovery of structure in the curriculum.These results suggest a promising new line of research for knowledge tracing. Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J. Guibas, Jascha Sohl-Dickstein |
NIPS | 1 |
| 2014 | Codewebs: scalable homework search for massive open online programming coursesabstractMassive open online courses (MOOCs), one of the latest internet revolutions have engendered hope that constant iterative improvement and economies of scale may cure the ``cost disease" of higher education. While scalable in many ways, providing feedback for homework submissions (particularly open-ended ones) remains a challenge in the online classroom. In courses where the student-teacher ratio can be ten thousand to one or worse, it is impossible for instructors to personally give feedback to students or to understand the multitude of student approaches and pitfalls. Organizing and making sense of massive collections of homework solutions is thus a critical web problem. Despite the challenges, the dense solution space sampling in highly structured homeworks for some MOOCs suggests an elegant solution to providing quality feedback to students on a massive scale. Andy Nguyen, Chris Piech, Jonathan Huang, Leonidas J. Guibas |
WWW | 2 |
| 2013 | Tuned Models of Peer Assessment in MOOCs
Chris Piech, Jonathan Huang, Chuong B. Do, Andrew Y. Ng, Daphne Koller |
EDM | 1 |
| 2013 | Deconstructing disengagement: analyzing learner subpopulations in massive open online coursesabstractAs MOOCs grow in popularity, the relatively low completion rates of learners has been a central criticism. This focus on completion rates, however, reflects a monolithic view of disengagement that does not allow MOOC designers to target interventions or develop adaptive course features for particular subpopulations of learners. To address this, we present a simple, scalable, and informative classification method that identifies a small number of longitudinal engagement trajectories in MOOCs. Learners are classified based on their patterns of interaction with video lectures and assessments, the primary features of most MOOCs to date. René F. Kizilcec, Chris Piech, Emily Schneider |
LAK | 2 |
| 2012 | Modeling how students learn to programabstractDespite the potential wealth of educational indicators expressed in a student's approach to homework assignments, how students arrive at their final solution is largely overlooked in university courses. In this paper we present a methodology which uses machine learning techniques to autonomously create a graphical model of how students in an introductory programming course progress through a homework assignment. We subsequently show that this model is predictive of which students will struggle with material presented later in the class. Chris Piech, Mehran Sahami, Daphne Koller, Steve Cooper, Paulo Blikstein |
SIGCSE | 1 |