VLDB 2026 Research / reviewers in the wild / expert
Juliette Woodrow
dblp:314/7435
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0006-8097-093XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aligning Small Language Models for Programming Feedback: Towards Scalable Coding Support in a Massive Global CourseabstractProviding timely and actionable feedback is essential for students learning to program. While large language models (LLMs) are increasingly used to automate this process, they remain costly to deploy and raise concerns around privacy and institutional control. Small language models (SLMs) offer a promising alternative: they can be run locally and integrated more flexibly into educational platforms. However, their out-of-the-box performance is often poor, requiring targeted training to be effective in classrooms. In this paper, we investigate whether a trained 3B-parameter SLM, guided by rubric-based prompting and a pipeline combining supervised and preference-based learning, can generate diagnostic feedback that approaches the quality of larger models. We deploy the model in a large-scale online programming course and compare its feedback to its base and fine-tuned variants, Llama-3.1-8B, and GPT-4.1, using human ratings from 53 teaching assistants and an automated LLM-as-a-judge analysis. Our results show that careful training narrows the feedback quality gap between an SLM and an LLM from over 80 to just 10 percentage points on key metrics. The trained SLM more rarely hallucinates errors, is often rated as helpful by educators, and only occasionally misses issues in student code. These findings suggest that small models can serve as practical and scalable targeted feedback solutions in large educational settings, while LLMs may remain necessary for more comprehensive diagnostic feedback. Charles Koutcheme, Juliette Woodrow, Chris Piech |
SIGCSE (1) | 2 |
| 2025 | Improving Generative AI Student Feedback: Direct Preference Optimization with Teachers in the Loop
Juliette Woodrow, Chris Piech, Oluwasanmi Koyejo |
EDM | 1 |
| 2025 | Soft Grades: A Calibrated and Accurate Method for Course-Grade Estimation that Expresses UncertaintyabstractIn traditional educational settings, students are often summarized by a single number-a final course grade-that reflects their performance.While final grades are convenient for reporting or comparison, they oversimplify a student's true ability and do not express uncertainty.In this paper, we introduce a new item-response model for classroom settings that infers a distribution over student abilities and uses this to represent each student's final grade as a probability distribution.This approach captures the uncertainty that comes from variations in both student performance and grading processes.Practical applications of our approach include enabling teachers to better understand grading confidence, impute missing assignment scores, and make informed decisions when curving final grades.For students, the model offers probabilistic estimates of their final course grades based on current performance, supporting informed academic decisions such as opting for Pass/Fail grading.We evaluate our model using real-world datasets, showing that the Soft Grades model is well-calibrated and surpasses the state-of-the-art polytomous IRT model in accurately predicting future scores.Additionally, we share a web application and Python scripts to make our model available to teachers and students. Juliette Woodrow, Chris Piech |
LAK | 1 |
| 2025 | The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam PerformancesabstractLarge language models (LLMs) are quickly being adopted in a wide range of learning experiences, especially via ubiquitous and broadly accessible chat interfaces like ChatGPT. This type of interface is readily available to students and teachers around the world. Coding education is an interesting test case, both because LLMs have strong performance on coding tasks, and because LLM-powered support tools are rapidly becoming part of the workflow of professional software engineers. To help understand the impact of generic LLM use on coding education, we conducted a large-scale randomized control trial with 5,831 students from 146 countries in an online coding class in which we provided some students with access to a chat interface with GPT-4. Under some assumptions, we estimate positive benefits on exam performance for adopters, the students who used the tool, but over all students, the advertisement of GPT-4 led to a significant average decrease in exam participation. We observe similar decreases in other forms of course engagement. However, this decrease is modulated by the student's country of origin. Offering access to LLMs to students from low human development index countries increased their exam participation rate on average. Our results suggest there may be promising benefits to using LLMs in an introductory coding class, but also potential harms for engagement, which makes their longer term impact on student success unclear. Our work highlights the need for additional investigations to help understand the potential impact of future adoption and integration of LLMs into classrooms. Allen Nie, Yash Chandak, Miroslav Suzara, Ali Malik, Juliette Woodrow, Matt Peng, Mehran Sahami, Emma Brunskill, Chris Piech |
L@S | 5 |
| 2025 | Infinite Story
Chris Piech, Mehran Sahami, Yasmine Alonso, Katie Liu, Javokhir Arifov, Anjali Sreenivas, Dan Webber, Tina Zheng, Ngoc Nguyen, Iddah Mlauzi, Juliette Woodrow |
SIGCSE (2) | 11 |
| 2024 | TeachNow: Enabling Teachers to Provide Spontaneous, Realtime 1: 1 Help in Massive Online CoursesabstractOne-on-one help from a teacher is highly impactful for students, yet extremely challenging to support in massive online courses (MOOCs). In this work, we present TeachNow: a novel system that lets volunteer teachers from anywhere in the world instantly provide 1:1 help sessions to students in MOOCs, without any scheduling or coordination overhead. TeachNow works by quickly finding an online student to help and putting them in a collaborative working session with the teacher. The spontaneous, on-demand nature of TeachNow gives teachers the flexibility to help whenever their schedule allows. Ali Malik, Juliette Woodrow, Chris Piech |
ITiCSE (1) | 2 |
| 2024 | A Fast and Accurate Machine Learning Autograder for the Breakout AssignmentabstractIn this paper, we detail the successful deployment of a machine learning autograder that significantly decreases the grading labor required in the Breakout computer science assignment. This assignment - which tasks students with programming a game consisting of a controllable paddle and a ball that bounces off the paddle to break bricks - is popular for engaging students with introductory computer science concepts, but creates a large grading burden. Due to the game's interactive nature, grading defies traditional unit tests and instead typically requires 8+ minutes of manually playing each student's game to search for bugs. This amounts to 45+ hours of grading in a standard course offering and prevents further widespread adoption of the assignment. Our autograder alleviates this burden by playing each student's game with a reinforcement learning agent and providing videos of discovered bugs to instructors. In an A/B test with manual grading, we find that our human-in-the-loop AI autograder reduces grading time by 44%, while slightly improving grading accuracy by 6%, ultimately saving roughly 30 hours over our deployment in two offerings of the assignment. Our results further suggest the practicality of grading other interactive assignments (e.g., other games or building websites) via similar machine learning techniques. Live demo at https://ezliu.github.io/breakoutgrader. Evan Zheran Liu, David Yuan 0001, Elyse Cornwall, Juliette Woodrow, Kaylee Burns, Allen Nie, Emma Brunskill, Chris Piech, Chelsea Finn |
SIGCSE (1) | 5 |
| 2024 | Learners Teaching Novices: An Uplifting Alternative AssessmentabstractWe propose and carry-out a novel method of formative assessment called Assessment via Teaching (AVT), in which learners demonstrate their understanding of CS1 topics by tutoring more novice students. AVT has powerful benefits over traditional forms of assessment: it is centered around service to others and is highly rewarding for the learners who teach. Moreover, teaching greatly improves the learners' own understanding of the material and has a huge positive impact on novices, who receive free 1:1 tutoring. Lastly, this form of assessment is naturally difficult to cheat---a critical property for assessments in the era of large-language models. We use AVT in a randomised control trial with learners in a CS1 course at an R1 university. The learners provide tutoring sessions to more novice students taking a lagged online version of the same course. We show that learners who do an AVT session before the course exam performed 20 to 30 percentage points better than the class average on several questions. Moreover, compared to students who did a practice exam, the AVT learners enjoyed their experience more and were twice as likely to study for their teaching session. We believe AVT is a scalable and uplifting method for formative assessment that could one day replace traditional exams. Ali Malik, Juliette Woodrow, Chris Piech |
SIGCSE (1) | 2 |
| 2024 | AI Teaches the Art of Elegant Coding: Timely, Fair, and Helpful Style Feedback in a Global CourseabstractTeaching students how to write code that is elegant, reusable, and comprehensible is a fundamental part of CS1 education. However, providing this "style feedback" in a timely manner has proven difficult to scale. In this paper, we present our experience deploying a novel, real-time style feedback tool in Code in Place, a large-scale online CS1 course. Our tool is based on the latest breakthroughs in large-language models (LLMs) and was carefully designed to be safe and helpful for students. We used our Real-Time Style Feedback tool (RTSF) in a class with over 8,000 diverse students from across the globe and ran a randomized control trial to understand its benefits. We show that students who received style feedback in real-time were five times more likely to view and engage with their feedback compared to students who received delayed feedback. Moreover, those who viewed feedback were more likely to make significant style-related edits to their code, with over 79% of these edits directly incorporating their feedback. We also discuss the practicality and dangers of LLM-based tools for feedback, investigating the quality of the feedback generated, LLM limitations, and techniques for consistency, standardization, and safeguarding against demographic bias, all of which are crucial for a tool utilized by students. Juliette Woodrow, Ali Malik, Chris Piech |
SIGCSE (1) | 1 |
| 2022 | Nifty AssignmentsabstractThe Nifty Assignments special session is about sharing the ideas and ready-to-use materials of successful assignments. Nick Parlante, Julie Zelenski, Eric Roberts 0001, Jed Rembold, Ben Stephenson, Jonathan Hudson, Stephanie Valentine, Juliette Woodrow, Kathleen Creel, Nick Bowman, L. Joshua Crotts, Andrew Matzureff, Mike Izbicki |
SIGCSE (2) | 8 |