David H. Smith IV

dblp:240/5344 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-6572-4347ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Evaluating AI Models for Autograding Explain in Plain English Questions: Challenges and Considerations
abstract
Code-reading ability has traditionally been under-emphasized in assessments as it is difficult to assess at scale. Prior research has shown that code-reading and code-writing are closely related skills; thus being able to assess and train code reading skills may be necessary for student learning. One way to assess code-reading ability is using Explain in Plain English (EiPE) questions, which ask students to describe what a piece of code does with natural language. Previous research deployed a binary (correct/incorrect) autograder using bigram models that performed comparably with human teaching assistants on student responses. With a dataset of 3,064 student responses from 17 EiPE questions, we investigated multiple autograders for EiPE questions. We evaluated methods as simple as logistic regression trained on bigram features, to more complicated Support Vector Machines (SVMs) trained on embeddings from Large Language Models (LLMs) to GPT-4. We found multiple useful autograders, most with accuracies in the \(86\!\!-\!\!88\%\) range, with different advantages. SVMs trained on LLM embeddings had the highest accuracy; few-shot chat completion with GPT-4 required minimal human effort; pipelines with multiple autograders for specific dimensions (what we call 3D autograders) can provide fine-grained feedback; and code generation with GPT-4 to leverage automatic code testing as a grading mechanism in exchange for slightly more lenient grading standards. While piloting these autograders in a non-major introductory Python course, students had largely similar views of all autograders, although they more often found the GPT-based grader and code-generation graders more helpful and liked the code-generation grader the most.
Maxwell Fowler, Chinedu Emeka, Binglin Chen, David H. Smith IV, Matthew West 0001, Craig B. Zilles
ACM Trans. Interact. Intell. Syst.4
2024 How Instructors Incorporate Generative AI into Teaching Computing
abstract
Generative AI (GenAI) has seen great advancements in the past two years and the conversation around adoption is increasing. Widely available GenAI tools are disrupting classroom practices as they can write and explain code with minimal student prompting. While most acknowledge that there is no way to stop students from using such tools, a consensus has yet to form on how students should use them if they choose to do so. At the same time, researchers have begun to introduce new pedagogical tools that integrate GenAI into computing curricula. These new tools offer students personalized help or attempt to teach prompting skills without undercutting code comprehension. This working group aims to detail the current landscape of education-focused GenAI tools and teaching approaches, present gaps where new tools or approaches could appear, identify good practice-examples, and provide a guide for instructors to utilize GenAI as they continue to adapt to this new era.
James Prather, Juho Leinonen 0001, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Virginia Pettit, Leo Porter 0001, Brent N. Reeves, Jaromír Savelka, David H. Smith IV, Sven Strickroth, Daniel Zingaro
ITiCSE (2)13
2024 CS1-LLM: Integrating LLMs into CS1 Instruction
abstract
The recent, widespread availability of Large Language Models (LLMs) like ChatGPT and GitHub Copilot may impact introductory programming courses (CS1) both in terms of what should be taught and how to teach it. Indeed, recent research has shown that LLMs are capable of solving the majority of the assignments and exams we previously used in CS1. In addition, professional software engineers are often using these tools, raising the question of whether we should be training our students in their use as well. This experience report describes a CS1 course at a large research-intensive university that fully embraces the use of LLMs from the beginning of the course. To incorporate the LLMs, the course was intentionally altered to reduce emphasis on syntax and writing code from scratch. Instead, the course now emphasizes skills needed to successfully produce software with an LLM. This includes explaining code, testing code, and decomposing large problems into small functions that are solvable by an LLM. In addition to frequent, formative assessments of these skills, students were given three large, open-ended projects in three separate domains (data science, image processing, and game design) that allowed them to showcase their creativity in topics of their choosing. In an end-of-term survey, students reported that they appreciated learning with the assistance of the LLM and that they interacted with the LLM in a variety of ways when writing code. We provide lessons learned for instructors who may wish to incorporate LLMs into their course.
Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, Leo Porter 0001
ITiCSE (1)3
2023 Generating Multiple Choice Questions for Computing Courses Using Large Language Models
abstract
Generating high-quality multiple-choice questions (MCQs) is a time-consuming activity that has led practitioners and researchers to develop community question banks and reuse the same questions from semester to semester. This results in generic MCQs which are not relevant to every course. Template-based methods for generating MCQs require less effort but are similarly limited. At the same time, advances in natural language processing have resulted in large language models (LLMs) that are capable of doing tasks previously reserved for people, such as generating code, code explanations, and programming assignments. In this paper, we investigate whether these generative capabilities of LLMs can be used to craft high-quality M CQs more efficiently, thereby enabling instructors to focus on personalizing MCQs to each course and the associated learning goals. We used two LLMs, GPT-3 and GPT-4, to generate isomorphic MCQs based on MCQs from the Canterbury Question Bank and an Introductory to Low-level C Programming Course. We evaluated the resulting MCQs to assess their ability to generate correct answers based on the question stem, a task that was previously not possible. Finally, we investigate whether there is a correlation between model performance and the discrimination score of the associated MCQ to understand whether low discrimination questions required the model to do more inference and therefore perform poorly. GPT-4 correctly generated the answer for 78.5% of MCQs based only on the question stem. This suggests that instructors could use these models to quickly draft quizzes, such as during a live class, to identify misconceptions in real-time. We also replicate previous findings that GPT-3 performs poorly on answering, or in our case generating, correct answers to MCQs. We also present cases we observed where LLMs struggled to produce correct answers. Finally, we discuss implications for computing education.
Andrew Tran, Kenneth Angelikas, Egi Rama, Chiku Okechukwu, David H. Smith IV, Stephen MacNeil
FIE5
2023 "\"I Don't Gamble To Make My Livelihood\": Understanding the Incentives For
abstract
Background: Prior work has primarily been concerned with identifying: (1) how Open Education Resources (OERs) can be used to increase the availability of educational materials, (2) what motivations are behind their adoption and usage in classrooms, and (3) what barriers impede said adoption. However, there is relatively little work investigating the motives and barriers to contribution in OER.
Maxwell Fowler, David H. Smith IV, Binglin Chen, Craig B. Zilles
ICER (1)2
2023 Useful Distractions? Investigating the Utility of Distractors in Parsons Problems
abstract
Parsons problems have become an increasingly popular pedagogical tool for teaching students in introductory programming courses. Since their inception these questions have included “distractor blocks” which are blocks of code that contain errors representing common misconceptions that students have. However, there have been limited investigations into how these distractors should be made, evaluated, and the subsequent utility of their inclusion on summative and formative assessments. In this line of work, I seek to provide a comprehensive evaluation of their role in summative assessments and their impact on learning in formative settings.
David H. Smith IV
ICER (2)1
2023 Investigating the Effects of Testing Frequency on Programming Performance and Students' Behavior
abstract
We conducted an across-semester quasi-experimental study that compared students' outcomes under frequent and infrequent testing regimens in an introductory computer science course. Students in the frequent testing (4 quizzes and 4 exams) semester outperformed the infrequent testing (1 midterm and 1 final exam) semester by 9.1 to 13.5 percentage points on code writing questions.
David H. Smith IV, Chinedu Emeka, Maxwell Fowler, Matthew West 0001, Craig B. Zilles
SIGCSE (1)1
2022 Are We Fair?: Quantifying Score Impacts of Computer Science Exams with Randomized Question Pools
abstract
With the increase of large enrollment courses and the growing need to offer online instruction, computer-based exams randomly generated from question pools have a clear benefit for computing courses. Such exams can be used at scale, scheduled asynchronously and/or online, and use versioning to make attempts at cheating less profitable. Despite these benefits, we want to ensure that the technique is not unfair to students, particularly when it comes to equivalent difficulty across exam versions.
Maxwell Fowler, David H. Smith IV, Chinedu Emeka, Matthew West 0001, Craig B. Zilles
SIGCSE (1)2
2021 Towards Modeling Student Engagement with Interactive Computing Textbooks: An Empirical Study
abstract
Interactive textbooks have great potential to increase student engagement with the course content which is critical to effective learning in computing education. Prior research on digital textbooks and interactive visualizations contributes to our understanding of student interactions with visualizations and modeling textbook knowledge concepts. However, research investigating student usage of interactive computing textbooks is still lacking. This study seeks to fill this gap by modeling student engagement with a Jupyter-notebook-based interactive textbook. Our findings suggest that students' active interactions with the presented interactive textbook, including changing, adding, and executing code in addition to manipulating visualizations, are significantly stronger in predicting student performance than conventional reading metrics. Our findings contribute to a deeper understanding of student interactions with interactive textbooks and provide guidance on the effective usage of said textbooks in computing education.
David H. Smith IV, Christopher D. Hundhausen, Filip Jagodzinski, Josh Myers-Dean, Kira Jaeger
SIGCSE1
2019 Investigating the Essential of Meaningful Automated Formative Feedback for Programming Assignments
abstract
This study investigated the essential of meaningful automated feedback for programming assignments. Three different types of feedback were tested, including (a) What's wrong- what test cases were testing and which failed, (b) Gap comparisons between expected and actual outputs, and (c) Hint hints on how to fix problems if test cases failed. 46 students taking a CS2 participated in this study. They were divided into three groups, and the feedback configurations for each group were different: (1) Group One -What's wrong, (2) Group Two -What's wrong + Gap, (3) Group Three -What's wrong+ Gap + Hint. This study found that simply knowing what failed did not help students sufficiently, and might stimulate system gaming behavior. Hints were not found to be impactful on student performance or their usage of automated feedback. Based on the findings, this study provides practical guidance on the design of automated feedback.
Jack P. Wilson, Camille Ottaway, Naitra Iriumi, Kai Arakawa, David H. Smith IV
VL/HCC6
2019 A Systematic Investigation of Replications in Computing Education Research
abstract
As the societal demands for application and knowledge in computer science (CS) increase, CS student enrollment keeps growing rapidly around the world. By continuously improving the efficacy of computing education and providing guidelines for learning and teaching practice, computing education research plays a vital role in addressing both educational and societal challenges that emerge from the growth of CS students. Given the significant role of computing education research, it is important to ensure the reliability of studies in this field. The extent to which studies can be replicated in a field is one of the most important standards for reliability. Different fields have paid increasing attention to the replication rates of their studies, but the replication rate of computing education was never systematically studied. To fill this gap, this study investigated the replication rate of computing education between 2009 and 2018. We examined 2,269 published studies from three major conferences and two major journals in computing education, and found that the overall replication rate of computing education was 2.38%. This study demonstrated the need for more replication studies in computing education and discussed how to encourage replication studies through research initiatives and policy making.
David H. Smith IV, Naitra Iriumi, Michail Tsikerdekis, Amy J. Ko
ACM Trans. Comput. Educ.2