VLDB 2026 Research / reviewers in the wild / expert
David H. Smith
dblp:97/11493
· DBLP profile ↗
23ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 21 · 9 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Coding in the CS Classroom: An Interview Study of Instructor PracticesabstractBackground and Context. Live coding is a pedagogical technique where an instructor writes code in front of their class. The literature on live coding consists primarily of experiments and case studies, and few researchers have sought to understand how instructors actually employ live coding in their classrooms. Daniel Manesh, Vee Pettit, Yan Chen 0033, David H. Smith |
ICER (1) | 5 |
| 2026 | How Students Focus their Studying when Offered Exam Re-takes
Lucas Flygare, David H. Smith, Geoffrey L. Herman, Maxwell Fowler, Craig B. Zilles |
ITiCSE (1) | 2 |
| 2026 | Accessible Keyboard Controls for Block Ordering Problems
Peter Fowles, Andreas Stefik, Brianna Blaser, David H. Smith, Seth Poulsen |
ITiCSE (1) | 4 |
| 2026 | Using the Potential of GenAI Tools for Accessibility
Natalie Kiesler, Bedour Alshaigy, Yasmine N. El-Glaly, Ilenia Fronza, Alex Gerdes, Earl W. Huff Jr., Sven Jacobs, Dominic Lohr, Raymond Pettit, Andreas Scholl, Sandra Schulz 0001, David H. Smith |
ITiCSE (2) | 12 |
| 2026 | Purplex: An Integrated Platform for Natural Language Programming and Code Comprehension Activities
David H. Smith, Paul Denny 0001, Kaitlin Riegel |
ITiCSE (2) | 1 |
| 2026 | Enabling Open Educational Resource Adoption through Integrated Sharing in PrairieLearnabstractThis paper introduces the PrairieLearn Question Sharing System (PQSS), which enables instructors to share question generators with other instructors, either as open educational resources or privately. PQSS is integrated into PrairieLearn, an open-source, problem-driven online learning platform. PQSS addresses a critical need for more open-source assessments by making it easier for instructors to share assessments and for instructors to use those assessments. Instructors often do not share questions due to the time it takes to publish them and the lack of recognition for their work. Because it is directly integrated into PrairieLearn, PQSS reduces the aforementioned friction of sharing and using shared questions, and we can report usage statistics to help question authors receive recognition for their work. In this paper, we share design and implementation details of the system, as well as experiences using it to share course content across courses and between universities. Seth Poulsen, Geoffrey L. Herman, Mariana Silva, Maxwell Fowler, David H. Smith, Leo Porter 0001, Nico Ritschel, Craig B. Zilles, Matthew West 0001 |
SIGCSE (1) | 5 |
| 2026 | Prompting through Decomposition: Evaluating the Efficacy of Problem Decomposition Diagrams for Code GenerationabstractWhen engaged in the initial design of a program, novice programmers and seasoned developers alike often sketch out---or, perhaps more famously, whiteboard---their ideas. However, with the introduction of natively multimodal Generative AI models, such diagrams may now function as a means of code generation in their own right. In this work, we perform an initial evaluation to understand how student-created decomposition diagrams can serve as prompts for code generation, with implications for teaching and assessing problem decomposition skills. David H. Smith, S. Moonwara A. Monisha, Annapurna Vadaparty, Leo Porter 0001, Daniel Zingaro |
SIGCSE (2) | 1 |
| 2025 | Counting the Trees in the Forest: Evaluating Prompt Segmentation for Classifying Code Comprehension LevelabstractReading and understanding code are fundamental skills for novice programmers, and especially important with the growing prevalence of AI-generated code and the need to evaluate its accuracy and reliability. ''Explain in Plain English'' questions are a widely used approach for assessing code comprehension, but providing automated feedback, particularly on comprehension levels, is a challenging task. This paper introduces a novel method for automatically assessing the comprehension level of responses to ''Explain in Plain English'' questions. Central to this is the ability to distinguish between two response types: multi-structural, where students describe the code line-by-line, and relational, where they explain the code's overall purpose. Using a Large Language Model (LLM) to segment both the student's description and the code, we aim to determine whether the student describes each line individually (many segments) or the code as a whole (fewer segments). We evaluate this approach's effectiveness by comparing segmentation results with human classifications, achieving substantial agreement. We conclude with how this approach, which we release as an open source Python package, could be used as a formative feedback mechanism. David H. Smith, Maxwell Fowler, Paul Denny 0001, Craig B. Zilles |
ITiCSE (1) | 1 |
| 2025 | ReDefining Code Comprehension: Function Naming as a Mechanism for Evaluating Code Comprehensionabstract''Explain in Plain English'' (EiPE) questions are widely used to assess code comprehension skills but are challenging to grade automatically. Recent approaches like Code Generation Based Grading (CGBG) leverage large language models (LLMs) to generate code from student explanations and validate its equivalence to the original code using unit tests. However, this approach does not differentiate between high-level, purpose-focused responses and low-level, implementation-focused ones, limiting its effectiveness in assessing comprehension level. We propose a modified approach where students generate function names, emphasizing the function's purpose over implementation details. We evaluate this method in an introductory programming course and analyze it using Item Response Theory (IRT) to assess the difficulty and discrimination of function naming exercises as exam items and to compare their alignment with traditional EiPE grading standards. We also publish this work as an open source Python package for auto-grading EiPE questions, providing a scalable solution for adoption. David H. Smith, Maxwell Fowler, Paul Denny 0001, Craig B. Zilles |
ITiCSE (1) | 1 |
| 2025 | Neurodiversity in Computing Education Research: A Systematic Literature ReviewabstractEnsuring equitable access to computing education for all students---including those with autism, dyslexia, or ADHD---is essential to developing a diverse and inclusive workforce. To understand the state of disability research in computing education, we conducted a systematic literature review of research on neurodiversity in computing education. Our search resulted in 1,943 total papers, which we filtered to 14 papers based on our inclusion criteria. Our mixed-methods approach analyzed research methods, participants, contribution types, and findings. The three main contribution types included empirical contributions based on user studies (57.1%), opinion contributions and position papers (50%), and survey contributions (21.4%). Interviews were the most common methodology (75% of empirical contributions). There were often inconsistencies in how research methods were described (e.g., number of participants and interview and survey materials). Our work shows that research on neurodivergence in computing education is still very preliminary. Most papers provided curricular recommendations that lacked empirical evidence to support those recommendations. Three areas of future work include investigating the impacts of active learning, increasing awareness and knowledge about neurodiverse students' experiences, and engaging neurodivergent students in the design of pedagogical materials and computing education research. Cynthia Zastudil, David H. Smith, Yusef Tohamy, Rayhona Nasimova, Gavin Montross, Stephen MacNeil |
ITiCSE (1) | 2 |
| 2025 | Exploring Student Reactions to LLM-Generated Feedback on Explain in Plain English ProblemsabstractCode reading and comprehension skills are essential for novices learning programming, and explain-in-plain-English tasks (EiPE) are a well-established approach for assessing these skills. However, manual grading of EiPE tasks is time-consuming and this has limited their use in practice. To address this, we explore an approach where students explain code samples to a large language model (LLM) which generates code based on their explanations. This generated code is then evaluated using test suites, and shown to students along with the test results. We are interested in understanding how automated formative feedback from an LLM guides students' subsequent prompts towards solving EiPE tasks. We analyzed 177 unique attempts on four EiPE exercises from 21 students, looking at what kinds of mistakes they made and how they fixed them. We found that when students made mistakes, they identified and corrected them using either a combination of the LLM-generated code and test case results, or they switched from describing the purpose of the code to describing the sample code line-by-line until the LLM-generated code exactly matched the obfuscated sample code. Our findings suggest both optimism and caution with the use of LLMs for unmonitored formative feedback. We identified false positive and negative cases, helpful variable naming, and clues of direct code recitation by students. For most students, this approach represents an efficient way to demonstrate and assess their code comprehension skills. However, we also found evidence of misconceptions being reinforced, suggesting the need for further work to identify and guide students more effectively. Chris Kerslake, Paul Denny 0001, David H. Smith, Juho Leinonen 0001, Stephen MacNeil, Andrew Luxton-Reilly, Brett A. Becker |
SIGCSE (1) | 3 |
| 2025 | Mining Hierarchies with Conviction: Constructing the CS1 Skill Hierarchy with Pairwise Comparisons over Skill DistributionsabstractIntroductory Programming courses teach multiple skills such as 1) explaining the purpose of code, 2) the ability to arrange lines of code in correct sequence, and 3) the ability to trace through the execution of a program, and 4) the ability to write code from scratch. Knowing if a programming skill is a prerequisite to another would assist instructors in organizing their materials such that students encounter and learn new topics using optimal skill sequences. In this study, we used the conviction measure from association rule mining to perform pair-wise comparisons of five skills: Write, Trace, Reverse trace, Sequence, and Explain code. We used the data from four exams with more than 600 participants in each exam from a public university in the United States, where students solved programming assignments of different skills for several programming topics. Our findings matched the previous finding that tracing is a prerequisite for students to learn to write code. But, contradicting the previous claims, our analysis suggested that writing code is a prerequisite skill to explaining code and that sequencing code is not a prerequisite to writing code. Our research can help instructors by systematically arranging the skills students exercise when encountering a new topic. Dip Kiran Pradhan Newar, Maxwell Fowler, David H. Smith, Seth Poulsen |
SIGCSE (2) | 3 |
| 2024 | Distractors Make You Pay Attention: Investigating the Learning Outcomes of Including Distractor Blocks in Parsons ProblemsabstractBackground: In CS1 courses, Parsons problems are a popular activity in which students are given blocks of code and asked to rearrange them into the correct order. Parsons problems often include incorrect blocks of code referred to as distractor blocks. Despite their widespread use, there have been few investigations into how distractor blocks impact student learning. Objectives: Our goals are to understand (1) the impact that including distractor blocks in Parsons problems has on learning and (2) the causality underlying that learning, if any. Methods: In this paper, we present the results of an explanatory sequential mixed methods study investigating the impact of distractor blocks on student learning. For the initial, quantitative stage, we use a randomized control trial to quantify the learning outcomes from practice with Parsons problems that include distractor blocks, as measured via post-test taken immediately after the practice activity and a retention test taken a week later. This study is followed by think-aloud interviews with 10 students practicing using a mix of Parsons problems that do and do not contain distractors to understand differences in how students approach those problems. Findings: Our findings show that students who practiced using Parsons problems that contained distractors performed 11 percentage points better on the immediate post-test (statistically significant) and 10 percentage points better on the retention test (approaching significance). The results of the think-aloud interviews indicate that grouping distractors with blocks of correct code causes students to more closely attend to the details of the code within those blocks. Implications: The results of this study indicate that distractors are essential when Parsons problems are used in a formative context. When they are not included, students may be able to successfully place blocks of code without attending to details of the code. This in turn limits their ability to learn new concepts or reinforce existing knowledge from those code blocks. David H. Smith, Seth Poulsen, Chinedu Emeka, Zihan Wu 0002, Carl Christopher Haynes-Magyar, Craig B. Zilles |
ICER (1) | 1 |
| 2024 | Explaining Code with a Purpose: An Integrated Approach for Developing Code Comprehension and Prompting SkillsabstractPublisher Copyright: © 2024 Owner/Author. Paul Denny 0001, David H. Smith, Maxwell Fowler, James Prather, Brett A. Becker, Juho Leinonen 0001 |
ITiCSE (1) | 2 |
| 2024 | Quickly Producing "Isomorphic" Exercises: Quantifying the Impact of Programming Question PermutationsabstractSmall, auto-gradable programming exercises provide a useful tool with which to assess students' programming skills in introductory computer science. To reduce the time needed to produce programming exercises of similar difficulty, previous research has applied a permutation strategy to existing questions. Prior work has left several open questions: is prior exposure to a question typically indicative of higher student performance? Are observed changes in difficulty due to the specific surface feature permutations applied? How is student performance impacted by the first version of a question to which they may be exposed? Maxwell Fowler, David H. Smith, Craig B. Zilles |
ITiCSE (1) | 2 |
| 2024 | Code Generation Based Grading: Evaluating an Auto-grading Mechanism for "Explain-in-Plain-English" QuestionsabstractComprehending and conveying the purpose of code is often cited as being a key learning objective within introductory programming courses. To address this objective, "Explain in Plain English'' questions, where students are shown a segment of code and asked to provide an abstract description of the code's purpose, have been adopted. However, given EiPE questions require a natural language response, they often require manual grading which is time-consuming for course staff and delays feedback for students. With the advent of large language models (LLMs) capable of generating code, responses to EiPE questions can be used to generate code segments, the correctness of which can then be easily verified using test cases. We refer to this approach as "Code Generation Based Grading'' (CGBG) and in this paper we explore its agreement with human graders using EiPE responses from past exams in an introductory programming course taught in Python. Overall, we find that all CGBG approaches achieve moderate agreement with human graders with the primary area of disagreement being its leniency with respect to low-level and line-by-line descriptions of code. David H. Smith, Craig B. Zilles |
ITiCSE (1) | 1 |
| 2024 | Evaluating Micro Parsons Problems as Exam QuestionsabstractParsons problems are a type of programming activity that present learners with blocks of existing code and requiring them to arrange those blocks to form a program rather than write the code from scratch. Micro Parsons problems extend this concept by having students assemble segments of code to form a single line of code rather than an entire program. Recent investigations into micro Parsons problems have primarily focused on supporting learners leaving open the question of micro Parsons efficacy as an exam item and how students perceive it when preparing for exams.To fill this gap, we included a variety of micro Parsons problems on four exams in an introductory programming course taught in Python. We use Item Response Theory to investigate the difficulty of the micro Parsons problems as well as the ability of the questions to differentiate between high and low ability students. We then compare these results to results for related questions where students are asked to write a single line of code from scratch. Finally, we conduct a thematic analysis of the survey responses to investigate how students' perceptions of micro Parsons both when practicing for exams and as they appear on exams. Zihan Wu 0002, David H. Smith |
ITiCSE (1) | 2 |
| 2024 | Prompting for Comprehension: Exploring the Intersection of Explain in Plain English Questions and Prompt WritingabstractLearning to program requires the development of a variety of skills including the ability to read, comprehend, and communicate the purpose of code. In the age of large language models (LLMs), where code can be generated automatically, developing these skills is more important than ever for novice programmers. The ability to write precise natural language descriptions of desired behavior is essential for eliciting code from an LLM, and the code that is generated must be understood in order to evaluate its correctness and suitability. In introductory computer science courses, a common question type used to develop and assess code comprehension skill is the 'Explain in Plain English' (EiPE) question. In these questions, students are shown a segment of code and asked to provide a natural language description of that code's purpose. The adoption of EiPE questions at scale has been hindered by: 1) the difficulty of automatically grading short answer responses and 2) the ability to provide effective and transparent feedback to students. To address these shortcomings, we explore and evaluate a grading approach where a student's EiPE response is used to generate code via an LLM, and that code is evaluated against test cases to determine if the description of the code was accurate. This provides a scalable approach to creating code comprehension questions and enables feedback both through the code generated from a student's description and the results of test cases run on that code. We evaluate students' success in completing these tasks, their use of the feedback provided by the system, and their perceptions of the activity. David H. Smith, Paul Denny 0001, Maxwell Fowler |
L@S | 1 |
| 2024 | Evaluating Large Language Model Code Generation as an Autograding Mechanism for "Explain in Plain English" QuestionsabstractThe ability of students to ''Explain in Plain English'' (EiPE) the purpose of code is a critical skill for students in introductory programming courses to develop. EiPE questions serve as both a mechanism for students to develop and demonstrate code comprehension skills. However, evaluating this skill has been challenging as manual grading is time consuming and not easily automated. The process of constructing a prompt for the purposes of code generation for a Large Language Model, such OpenAI's GPT-4, bears a striking resemblance to constructing EiPE responses. In this paper, we explore the potential of using test cases run on code generated by GPT-4 from students' EiPE responses as a grading mechanism for EiPE questions. We applied this proposed grading method to a corpus of EiPE responses collected from past exams, then measured agreement between the results of this grading method and human graders. Overall, we find moderate agreement between the human raters and the results of the unit tests run on the generated code. This appears to be attributable to GPT-4's code generation being more lenient than human graders on low-level descriptions of code. David H. Smith, Craig B. Zilles |
SIGCSE (2) | 1 |
| 2023 | Conducting Multi-Institutional Studies of Parsons ProblemsabstractMany novice programmers struggle to write code from scratch and get frustrated when their code does not work. Parsons problems can reduce the difficulty of a coding problem by providing mixed-up blocks that the learner assembles in the correct order. Parsons problems can also include distractor blocks that are not needed in a correct solution, but which may help students learn to recognize and fix errors. Evidence indicates that students find Parsons problems engaging, easier than writing code from scratch, useful for learning patterns, and typically faster to solve than writing code from scratch with equivalent learning gains. This working group leverages the work of the 2022 ITiCSE working group which published an extensive literature review of Parsons problems and designed and piloted several studies based on the gaps identified by the literature review. The 2023 working group is revising, conducting, and creating new studies. We will analyze the data from these multi-institutional and multi-national studies and publish the results as well as recommendations for future working groups. Barbara Ericson, Janice L. Pearce, Susan H. Rodger, Andrew Csizmadia, Rita Garcia, Francisco J. Gutierrez, Konstantinos Liaskos, Aadarsh Padiyath, Michael 'Adrir' Scott, David H. Smith, Jayakrishnan Madathil Warriem, Angela M. Zavaleta Bernuy |
ITiCSE (2) | 10 |
| 2023 | Investigating the Role and Impact of Distractors on Parsons Problems in CS1 AssessmentsabstractIn recent years Parsons problems have grown in popularity as both a pedagogical tool and as an assessment item alike. In these problems, students are expected to take existing but jumbled blocks of code and organize them to form a working solution. It is common for these problems to include incorrect blocks of code, typically referred to as "distractors," alongside the correct blocks. However, the utility of these distractors and their impact on a problems difficulty has yet to be thoroughly investigated. This study contributes to filling this gap by comparing performance, time spent, and item discrimination statistics for 32 pairs of Parsons problems from CS1 Python exams and quizzes. Our findings indicate that the inclusion of distractors has a large impact on the amount of time students spend on the questions and a low to moderate impact on score. Additionally, problems without distractors were already found to have high discrimination and including distractors did little to improve their discrimination. These findings suggest that the inclusion of distractors does little to improve the quality of these problems as exam questions but may have a negative impact on students by causing them to spend significantly more time on the problems and reducing the time they have for the rest of the exam. David H. Smith, Maxwell Fowler, Craig B. Zilles |
ITiCSE (1) | 1 |
| 2023 | Discovering, Autogenerating, and Evaluating Distractors for Python Parsons Problems in CS1abstractIn this paper, we make three contributions related to the selection and use of distractors (lines of code reflecting common errors or misconceptions) in Parsons problems. First, we demonstrate a process by which templates for creating distractors can be selected through the analysis of student submissions to short answer questions. Second, we describe the creation of a tool that uses these templates to automatically generate distractors for novel problems. Third, we perform a preliminary analysis of how the presence of distractors impacts performance, problem solving efficiency, and item discrimination when used in summative assessments. Our results suggest that distractors should not be used in summative assessments because they significantly increase the problem's completion time without a significant increase in problem discrimination. David H. Smith, Craig B. Zilles |
SIGCSE (1) | 1 |
| 2017 | Computerized prescriber order entry-related patient safety reports: analysis of 2522 medication errorsabstractObjective: To examine medication errors potentially related to computerized prescriber order entry (CPOE) and refine a previously published taxonomy to classify them. Materials and Methods: We reviewed all patient safety medication reports that occurred in the medication ordering phase from 6 sites participating in a United States Food and Drug Administration-sponsored project examining CPOE safety. Two pharmacists independently reviewed each report to confirm whether the error occurred in the ordering/prescribing phase and was related to CPOE. For those related to CPOE, we assessed whether CPOE facilitated (actively contributed to) the error or failed to prevent the error (did not directly cause it, but optimal systems could have potentially prevented it). A previously developed taxonomy was iteratively refined to classify the reports. Results: Of 2522 medication error reports, 1308 (51.9%) were related to CPOE. Of these, CPOE facilitated the error in 171 (13.1%) and potentially could have prevented the error in 1137 (86.9%). The most frequent categories of "what happened to the patient" were delays in medication reaching the patient, potentially receiving duplicate drugs, or receiving a higher dose than indicated. The most frequent categories for "what happened in CPOE" included orders not routed to or received at the intended location, wrong dose ordered, and duplicate orders. Variations were seen in the format, categorization, and quality of reports, resulting in error causation being assignable in only 403 instances (31%). Discussion and Conclusion: Errors related to CPOE commonly involved transmission errors, erroneous dosing, and duplicate orders. More standardized safety reporting using a common taxonomy could help health care systems and vendors learn and implement prevention strategies. Mary G. Amato, Alejandra Salazar, Thu-Trang T. Hickman, Arbor J. L. Quist, Lynn A. Volk, Adam Wright, Dustin McEvoy, William L. Galanter, Ross Koppel, Beverly Loudin, Jason S. Adelman, John D. McGreevey, David H. Smith, David W. Bates, Gordon D. Schiff |
J. Am. Medical Informatics Assoc. | 13 |