John DeNero

dblp:91/4310 · DBLP profile ↗
← Back
49ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0001-9152-3891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 8 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 15 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Systems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
Rose Niousha, Samantha Boatright Smith, Bita Akram, Peter Brusilovsky, Arto Hellas, Juho Leinonen 0001, John DeNero, Narges Norouzi
AIED7
2026 Instructors' Perspectives on LLM-Generated Programming Formative Feedback
abstract
We study instructor perspectives on LLM-generated programming feedback in an introductory Python course. LLM tutors predominantly offered debugging help, while human instructors preferred more diverse feedback types, including conceptual reminders, revisiting the problem, and examples. Cases where LLM tutor feedback diverged from human instructors' intent required major edits with different feedback types, while cases with closer alignment needed only minor changes with similar feedback types. Findings highlight the need for LLM tutors to reflect on instructor intent to ensure pedagogically aligned feedback.
Rose Niousha, Samantha Boatright Smith, Abigail O'Neill, J. D. Zamfirescu-Pereira, John DeNero, Narges Norouzi
SIGCSE (2)5
2026 Misconception-Aware LLM Programming Tutor: Lessons Learned from Student-Tutor Interactions
abstract
Large Language Models (LLMs) are increasingly used as programming tutors, but their feedback is often generic and prone to solution leakage. To address these issues, we present MisconceptionTutor, which grounds feedback in common student misconceptions. Through both pre-deployment analyses and a real-classroom deployment, we find that even simple prompting frameworks can meaningfully steer tutor behavior to be more pedagogically oriented and noticeably more satisfying to students.
Rose Niousha, Samantha Boatright Smith, Abigail O'Neill, J. D. Zamfirescu-Pereira, John DeNero, Narges Norouzi
SIGCSE (2)5
2025 From Code to Concepts: Textbook-Driven Knowledge Tracing with LLMs in CS1
abstract
Gauging a student's understanding of course concepts, at an arbitrary point during a course, can be challenging. Standardized exams offer only a snapshot of performance rather than a deep understanding of progress. However, with Large Language Models (LLMs) now deployed at scale in CS1 courses, we can track multiple attempts from each student for every homework problem. This data provides insights into how students learn and deploy concepts over time, presenting a unique opportunity to rethink how we track changes in individual student knowledge. Traditional Knowledge Tracing (KT) methods often lack explainability and are computationally expensive. In contrast, our framework leverages an LLM to identify student progress on labeled, problem-level concepts from a student homework code submission. Our initial results show that the student's knowledge state can be dynamically updated. This knowledge state can then be used to provide more targeted, effective feedback and create tailored study materials.
Abigail O'Neill, Samantha Boatright Smith, Aneesh Durai, John DeNero, J. D. Zamfirescu-Pereira, Narges Norouzi
SIGCSE (2)4
2025 Spotting AI Missteps: Students Take on LLM Errors in CS1
Samantha Boatright Smith, Heather Wei, Abigail O'Neill, Aneesh Durai, John DeNero, J. D. Zamfirescu-Pereira, Narges Norouzi
SIGCSE (2)5
2025 Pensieve Discuss: Scalable Small-Group CS Tutoring System with AI
Yoonseok Yang, Jack Liu, J. D. Zamfirescu-Pereira, John DeNero
SIGCSE (2)4
2025 61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?)
abstract
LLM-based chatbots enable students to get immediate, interactive help on homework assignments, but even a thoughtfully-designed bot may not serve all pedagogical goals. We report here on the development and deployment of a GPT-4-based interactive homework assistant ("61A Bot'') for students in a large CS1 course; over 2000 students made over 100,000 requests of our Bot across two semesters. Our assistant offers one-shot, contextual feedback within the command-line "autograder'' students use to test their code. Our Bot wraps student code in a custom prompt that supports our pedagogical goals and avoids providing solutions directly. Analyzing student feedback, questions, and autograder data, we find reductions in homework-related question rates in our course forum, as well as reductions in homework completion time when our Bot is available. For students in the 50th -80th percentile, reductions can exceed 30 minutes per assignment, up to 50% less time than students at the same percentile rank in prior semesters. Finally, we discuss these observations, potential impacts on student learning, and other potential costs and benefits of AI assistance in CS1.
J. D. Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, Narges Norouzi
SIGCSE (1)4
2022 Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia
abstract
While neural networks demonstrate a remarkable ability to model linguistic content, capturing contextual information related to a speaker's conversational role is an open area of research.In this work, we analyze the effect of speaker role on language use through the game of Mafia, in which participants are assigned either an honest or a deceptive role.In addition to building a framework to collect a dataset of Mafia game records, we demonstrate that there are differences in the language produced by players with different roles.We confirm that classification models are able to rank deceptive players as more suspicious than honest ones based only on their use of language.Furthermore, we show that training models on two auxiliary tasks outperforms a standard BERT-based text classification approach.We also present methods for using our trained models to identify features that distinguish between player roles, which could be used to assist players during the Mafia game.
Samee Ibraheem, Gaoyue Zhou, John DeNero
NAACL-HLT3
2022 Automatic Correction of Human Translations
abstract
Jessy Lin, Geza Kovacs, Aditya Shastry, Joern Wuebker, John DeNero. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jessy Lin, Geza Kovacs, Jörn Wübker, John DeNero
NAACL-HLT5
2022 Innovation in Data Science Education
abstract
The workshop will allow participants to gain experience with a series of innovations developed at UC Berkeley that have enabled the teaching of undergraduate data science at scale to students from all backgrounds. Rather than beginning with established introductory strategies as the gateway to computer science, students in the "Foundations of Data Science" (data8.org) learn computational skills and concepts in relation to real-world issues and with attention to societal implications. By engaging with students' interest in the applications of computing on data, and integrating societal impact from the start, the program has developed a long-term commitment to advance computational skills for large numbers of students. These innovations in teaching not only convey important computational content, but also broaden participation beyond existing approaches to computer science. Goals include increasing diversity among students learning computer science, giving students a strong ethical foundation within their computer science work, and encouraging critical thinking in the application of inference and statistical techniques. Bringing a laptop is recommended. UC Berkeley has as over 1000 students in a large and open Data Science Major, where a range of Domain Emphases and backgrounds bring a broader set of students than the traditional CS major. UC Berkeley Data Science has been gathering over 500 educators in a summer workshop on sharing curricular innovation.
Eric Van Dusen, John DeNero, Kseniya Usovich
SIGCSE (2)2
2021 Supervised line attention for tumor attribute classification from pathology reports: Higher performance with less data
Nick Altieri 0001, Briton Park, Mara Olson, John DeNero, Anobel Y. Odisho, Bin Yu 0001
J. Biomed. Informatics4
2020 End-to-End Neural Word Alignment Outperforms GIZA++
abstract
Word alignment was once a core unsupervised learning task in natural language processing because of its essential role in training statistical machine translation (MT) models.Although unnecessary for training neural MT models, word alignment still plays an important role in interactive applications of neural machine translation, such as annotation transfer and lexicon injection.While statistical MT methods have been replaced by neural approaches with superior performance, the twenty-year-old GIZA++ toolkit remains a key component of state-of-the-art word alignment systems.Prior work on neural word alignment has only been able to outperform GIZA++ by using its output during training.We present the first end-to-end neural word alignment method that consistently outperforms GIZA++ on three data sets.Our approach repurposes a Transformer model trained for supervised translation to also serve as an unsupervised word alignment model in a manner that is tightly integrated and does not affect translation quality.
Thomas Zenkel, Jörn Wübker, John DeNero
ACL3
2020 Investigating the Behavior of Malicious Actors Through the Game of Mafia
Samee Ibraheem, Vael Gates, John DeNero, Thomas L. Griffiths 0001
CogSci3
2020 A Streaming Approach For Efficient Batched Beam Search
abstract
We propose an efficient batching strategy for variable-length decoding on GPU architectures.During decoding, when candidates terminate or are pruned according to heuristics, our streaming approach periodically "refills" the batch before proceeding with a selected subset of candidates.We apply our method to variable-width beam search on a state-of-theart machine translation model.Our method decreases runtime by up to 71% compared to a fixed-width beam search baseline and 17% compared to a variable-width baseline, while matching baselines' BLEU.Finally, experiments show that our method can speed up decoding in other domains, such as semantic and syntactic parsing.
Kevin Yang, Violet Yao, John DeNero, Daniel Klein 0001
EMNLP (1)3
2020 Nifty Assignments
abstract
The Nifty Assignments special session is about promoting and sharing the ideas and ready-to-use materials of successful assignments. Each presenter will introduce their assignment, give a quick demo, and describe its niche in the curriculum and its strengths and weaknesses. The presentations (and the descriptions below) merely introduce the assignment. A key part of Nifty Assignments is the mundane but vital role of distributing the materials - handouts, data files, starter code, rubrics, autograders - that make each assignment ready to adopt. Each assignment presented has complete materials freely available on the Nifty Assignments home page nifty.stanford.edu. If you have an assignment that works well and would be of interest to the CSE community, please consider applying to present at Nifty Assignments.
Nick Parlante, Julie Zelenski, John DeNero, Christopher Allsman, Tiffany Perumpail, Rahul Arya, Kavi Gupta, Catherine Cang, Paul Bitutsky, Ryan Moughan, David J. Malan, Brian Yu, Evan M. Peck, Carl Albing, Kevin Wayne, Keith Schwarz
SIGCSE3
2019 Guiding Policies with Language via Meta-Learning
John D. Co-Reyes, Abhishek Gupta 0004, Suvansh Sanjeev, Nick Altieri 0001, Jacob Andreas, John DeNero, Pieter Abbeel, Sergey Levine
ICLR (Poster)6
2018 Compact Personalized Models for Neural Machine Translation
abstract
We propose and compare methods for gradientbased domain adaptation of self-attentive neural machine translation models.We demonstrate that a large proportion of model parameters can be frozen during adaptation with minimal or no reduction in translation quality by encouraging structured sparsity in the set of offset tensors during learning via group lasso regularization.We evaluate this technique for both batch and incremental adaptation across multiple data sets and language pairs.Our system architecture-combining a state-of-the-art self-attentive model with compact domain adaptation-provides high quality personalized machine translation that is both space and time efficient.
Jörn Wübker, Patrick Simianer, John DeNero
EMNLP3
2017 Learning an Interactive Attention Policy for Neural Machine Translation
Samee Ibraheem, Nick Altieri 0001, John DeNero
MTSummit (1)3
2017 Beyond Autograding: Advances in Student Feedback Platforms
abstract
No abstract available.
John DeNero, Sumukh Sridhara, Manuel A. Pérez-Quiñones, Aatish Nayak, Ben Leong
SIGCSE1
2016 Models and Inference for Prefix-Constrained Machine Translation
abstract
We apply phrase-based and neural models to a core task in interactive machine translation: suggesting how to complete a partial translation.For the phrase-based system, we demonstrate improvements in suggestion quality using novel objective functions, learning techniques, and inference algorithms tailored to this task.Our contributions include new tunable metrics, an improved beam search strategy, an n-best extraction method that increases suggestion diversity, and a tuning procedure for a hierarchical joint model of alignment and translation.The combination of these techniques improves next-word suggestion accuracy dramatically from 28.5% to 41.2% in a large-scale English-German experiment.Our recurrent neural translation system increases accuracy yet further to 53.0%, but inference is two orders of magnitude slower.Manual error analysis shows the strengths and weaknesses of both approaches.
Jörn Wübker, Spence Green, John DeNero, Sasa Hasan, Minh-Thang Luong
ACL (1)3
2016 An Analysis of the Ability of Statistical Language Models to Capture the Structural Properties of Language
abstract
We investigate the characteristics and quantifiable predispositions of both n-gram and recurrent neural language models in the framework of language generation.In modern applications, neural models have been widely adopted, as they have empirically provided better results.However, there is a lack of deep analysis of the models and how they relate to real language and its structural properties.We attempt to perform such an investigation by analyzing corpora generated by sampling from the models.The results are compared to each other and to the results of the same analysis applied to the training corpus.We carried out these experiments on varieties of Kneser-Ney smoothed n-gram models and basic recurrent neural language models.Our results reveal a number of distinctive characteristics of each model, and offer insights into their behavior.Our general approach also provides a framework in which to perform further analysis of language models.
Aneiss Ghodsi, John DeNero
INLG2
2016 Fuzz Testing Projects in Massive Courses
abstract
Scaffolded projects with automated feedback are core instructional components of many massive courses. In subjects that include programming, feedback is typically provided by test cases constructed manually by the instructor. This paper explores the effectiveness of fuzz testing, a randomized technique for verifying the behavior of programs. In particular, we apply fuzz testing to identify when a student's solution differs in behavior from a reference implementation by randomly exploring the space of legal inputs to a program. Fuzz testing serves as a useful complement to manually constructed tests. Instructors can concentrate on designing targeted tests that focus attention on specific issues while using fuzz testing for comprehensive error checking. In the first project of a 1,400-student introductory computer science course, fuzz testing caught errors that were missed by a suite of targeted test cases for more than 48% of students. As a result, the students dedicated substantially more effort to mastering the nuances of the assignment.
Sumukh Sridhara, Brian Hou, Jeffrey Lu, John DeNero
L@S4
2016 Identifying Student Misunderstandings using Constructed Responses
abstract
In contrast to multiple-choice or selected response questions, constructed response questions can result in a wide variety of incorrect responses. However, constructed responses are richer in information. We propose a technique for using each student's constructed responses in order to identify a subset of their stable conceptual misunderstandings. Our approach is designed for courses with so many students that it is infeasible to interpret every distinct wrong answer manually. Instead, we label only the most frequent wrong answers with the misunderstandings that they indicate, then predict the misunderstandings associated with other wrong answers using statistical co-occurrence patterns. This tiered approach leverages a small amount of human labeling effort to seed an automated procedure that identifies misunderstandings in students. Our approach involves much less effort than inspecting all answers, substantially outperforms a baseline that does not take advantage of co-occurrence statistics, proves robust to different course sizes, and generalizes effectively across student cohorts.
Kristin Stephens-Martinez, An Ju, Colin Schoen, John DeNero, Armando Fox
L@S4
2016 CS10K Teachers by 2017?: Try CS1K+ students NOW! Coping with the Largest CS1 Courses in History
abstract
"Be careful what you wish for, you just might get it." - Proverb In 2005, computing education was experiencing a crisis. Enrollments had "fallen to such an extent that some academic computing programs were facing significant reductions in staffing levels or even elimination". The community responded, with panels to investigate and highlight ways to infuse "passion, beauty, joy and awe" into the introductory experiences, the CS10K project to bring computing to 10,000 teachers and 100,000 students, and better messaging of career opportunities, to name a few of the initiatives to bring students back into our seats.
Dan Garcia 0001, Jennifer Campbell, John DeNero, Mary Lou Dorf, Stuart Reges
SIGCSE3
2016 Nifty Assignments
abstract
I suspect that students learn more from our programming assignments than from our much sweated-over lectures, with their slide transitions, clip art, and joke attempts. A great assignment is deliberate about where the student hours go, concentrating the student's attention on material that is interesting and useful. The best assignments solve a problem that is topical and entertaining, providing motivation for the whole stack of work. Unfortunately, creating great programming assignments is both time consuming and error prone. The Nifty Assignments special session is all about promoting and sharing the ideas and ready-to-use materials of successful assignments.
Nick Parlante, Julie Zelenski, Baker Franke, Arvind Bhusnurmath, Karen Her, Kristen Gee, Eric D. Manley, Timothy Urness, Marvin Zhang, Brian Hou, John DeNero, Josh Hug, Kevin Wayne
SIGCSE11
2015 Variable-Length Word Encodings for Neural Translation Models
abstract
Recent work in neural machine translation has shown promising performance, but the most effective architectures do not scale naturally to large vocabulary sizes.We propose and compare three variable-length encoding schemes that represent a large vocabulary corpus using a much smaller vocabulary with no loss in information.Common words are unaffected by our encoding, but rare words are encoded using a sequence of two pseudo-words.Our method is simple and effective: it requires no complete dictionaries, learning procedures, increased training time, changes to the model, or new parameters.Compared to a baseline that replaces all rare words with an unknown word symbol, our best variable-length encoding strategy improves WMT English-French translation performance by up to 1.7 BLEU.
Rohan Chitnis, John DeNero
EMNLP2
2015 Hierarchical Incremental Adaptation for Statistical Machine Translation
abstract
We present an incremental adaptation approach for statistical machine translation that maintains a flexible hierarchical domain structure within a single consistent model.Both weights and rules are updated incrementally on a stream of post-edits.Our multi-level domain hierarchy allows the system to adapt simultaneously towards local context at different levels of granularity, including genres and individual documents.Our experiments show consistent improvements in translation quality from all components of our approach.
Jörn Wübker, Spence Green, John DeNero
EMNLP3
2015 Problems Before Solutions: Automated Problem Clarification at Scale
abstract
Automatic assessment reduces the need for individual feedback in massive courses, but often focuses only on scoring solutions, rather than assessing whether students correctly understand problems. We present an enriched approach to automatic assessment that explicitly assists students in understanding the detailed specification of technical problems that they are asked to solve, in addition to evaluating their solutions. Students are given a suite of solution test cases, but they must first unlock each test case by validating its behavior before they are allowed to apply it to their proposed solution. When provided with this automated feedback early in the problem-solving process, students ask fewer clarificatory questions and express less confusion about assessments. As a result, instructors spend less time explaining problems to students. In a 1300-person university course, we observed that the vast majority of students chose to validate their understanding of test cases before attempting to solve problems. These students reported that the validation process improved their understanding.
Soumya Basu 0003, Albert Wu, Brian Hou, John DeNero
L@S4
2014 A Constrained Viterbi Relaxation for Bidirectional Word Alignment
abstract
Bidirectional models of word alignment are an appealing alternative to post-hoc combinations of directional word aligners.Unfortunately, most bidirectional formulations are NP-Hard to solve, and a previous attempt to use a relaxationbased decoder yielded few exact solutions (6%).We present a novel relaxation for decoding the bidirectional model of DeNero and Macherey (2011).The relaxation can be solved with a modified version of the Viterbi algorithm.To find optimal solutions on difficult instances, we alternate between incrementally adding constraints and applying optimality-preserving coarse-to-fine pruning.The algorithm finds provably exact solutions on 86% of sentence pairs and shows improvements over directional models.
Yin-Wen Chang, Alexander M. Rush, John DeNero, Michael Collins 0001
ACL (1)3
2014 Teaching composition quality at scale: human judgment in the age of autograders
abstract
We describe an effort to improve the composition quality of student programs: the property that a program can be understood effectively by another person. As a semester-long component of UC Berkeley's first course for majors, CS 61A, we gave students composition guidelines, scores, and qualitative feedback-all generated manually by a course staff of 10 graders for over 700 students. To facilitate this effort, we created a new online tool that allows instructors to provide feedback efficiently at scale. Our system differs from recently developed alternatives in that it is a branch of an industrial tool originally developed for internal code reviews at Google and used extensively by the open-source community. We found that many of the features designed for industrial applications are well-suited for instructional use as well. We extended the system with permissions controls and comment memories tailored for giving educational feedback.
John DeNero, Stephen Martinis
SIGCSE1
2014 Nifty assignments
abstract
A great CS assignment is a delight to all, but the path to one can be most roundabout. Many CS students have had their characters built up on assignments that worked better as an idea than as an actual assignment. Assignments are hard to come up with, yet they are the key to student learning. The Nifty Assignments special session is all about promoting and sharing the ideas and ready-to-use materials of successful assignments.
Nick Parlante, Julie Zelenski, Josh Hug, John DeNero, Antti Laaksonen, Arto Vihavainen, Frank McCown, Kevin Wayne
SIGCSE5
2013 Identifying Phrasal Verbs Using Many Bilingual Corpora
abstract
We address the problem of identifying multiword expressions in a language, focusing on English phrasal verbs.Our polyglot ranking approach integrates frequency statistics from translated corpora in 50 different languages.Our experimental evaluation demonstrates that combining statistical evidence from many parallel corpora using a novel ranking-oriented boosting algorithm produces a comprehensive set of English phrasal verbs, achieving performance comparable to a human-curated set. *Research conducted during an internship at Google.
Karl Pichotta, John DeNero
EMNLP2
2013 Supervised Learning of Complete Morphological Paradigms
Greg Durrett, John DeNero
HLT-NAACL2
2013 Nifty assignments
abstract
Every time I re-use a handout, I look it over and make a few little "improvements". I play around with code demos and entertain myself with different slide transitions. However, inevitably, I return to the conclusion that most of what my students learn in my course comes from the assignments. Great assignments are hard to dream up and time-consuming to develop. With that in mind, the Nifty Assignments session is all about promoting and sharing the ideas and ready-to-use materials of successful assignments.
Nick Parlante, Julie Zelenski, Michelle Craig, John DeNero, Mark Guzdial, David J. Malan, Aditi S. Muralidharan, Eric Roberts 0001, Kevin Wayne
SIGCSE4
2012 A Class-Based Agreement Model for Generating Accurately Inflected Translations
Spence Green, John DeNero
ACL (1)2
2012 Unsupervised Translation Sense Clustering
Mohit Bansal, John DeNero, Dekang Lin
HLT-NAACL2
2011 Model-Based Aligner Combination Using Dual Decomposition
John DeNero, Klaus Macherey
ACL1
2011 Inducing Sentence Structure from Parallel Corpora for Reordering
John DeNero, Jakob Uszkoreit
EMNLP1
2010 Discriminative Modeling of Extraction Sets for Machine Translation
John DeNero, Daniel Klein 0001
ACL1
2010 Painless Unsupervised Learning with Features
Taylor Berg-Kirkpatrick, Alexandre Bouchard-Côté, John DeNero, Daniel Klein 0001
HLT-NAACL3
2010 Model Combination for Machine Translation
John DeNero, Shankar Kumar, Ciprian Chelba, Franz Josef Och
HLT-NAACL1
2009 Fast Consensus Decoding over Translation Forests
John DeNero, David Chiang 0001, Kevin Knight
ACL/IJCNLP1
2009 Better Word Alignments with Supervised ITG Models
Aria Haghighi, John Blitzer, John DeNero, Daniel Klein 0001
ACL/IJCNLP3
2009 Consensus Training for Consensus Decoding in Machine Translation
Adam Pauls, John DeNero, Daniel Klein 0001
EMNLP2
2009 Efficient Parsing for Transducer Grammars
John DeNero, Mohit Bansal, Adam Pauls, Daniel Klein 0001
HLT-NAACL1
2008 Sampling Alignment Structure under a Bayesian Translation Model
John DeNero, Alexandre Bouchard-Côté, Daniel Klein 0001
EMNLP1
2007 A* Search via Approximate Factoring
Aria Haghighi, John DeNero, Daniel Klein 0001
AAAI2
2007 Tailoring Word Alignments to Syntactic Machine Translation
John DeNero, Daniel Klein 0001
ACL1
2007 Approximate Factoring for A* Search
Aria Haghighi, John DeNero, Daniel Klein 0001
HLT-NAACL2