EDBT 2026 Demo / reviewers in the wild / expert
Matthew West 0001
dblp:16/1726-1 · also Matt West 0001
· DBLP profile ↗
57ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0002-7605-0050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 34 · 24 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 5 since 2021Systems, architecture and hardware · 10 · 1 first-author · 4 since 2021Theory of computation · 4Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
Jacob Levine, Miguel Aenlle, Craig B. Zilles, Matthew West 0001, Mariana Silva |
AIED | 4 |
| 2026 | A Two-Stage LLM Pipeline for Handwritten Mathematics AutogradingabstractWhile question-asking platforms have provided a wide array of pedagogical benefits and have decreased grader workload in university classes, they are limited in what types of inputs are accepted. Recent work has explored using Large Language Models to expand what can be automatically graded. In this study, we examine their use for grading handwritten mathematics questions via rubrics. Although much work remains, our preliminary results suggest that using separate prompts to first extract text and then evaluate rubric items enables LLMs to distinguish fully correct solutions from those requiring further review, thereby reducing grader workload. Jacob Levine, Matthew West 0001, Mariana Silva |
SIGCSE (2) | 2 |
| 2026 | Enabling Open Educational Resource Adoption through Integrated Sharing in PrairieLearnabstractThis paper introduces the PrairieLearn Question Sharing System (PQSS), which enables instructors to share question generators with other instructors, either as open educational resources or privately. PQSS is integrated into PrairieLearn, an open-source, problem-driven online learning platform. PQSS addresses a critical need for more open-source assessments by making it easier for instructors to share assessments and for instructors to use those assessments. Instructors often do not share questions due to the time it takes to publish them and the lack of recognition for their work. Because it is directly integrated into PrairieLearn, PQSS reduces the aforementioned friction of sharing and using shared questions, and we can report usage statistics to help question authors receive recognition for their work. In this paper, we share design and implementation details of the system, as well as experiences using it to share course content across courses and between universities. Seth Poulsen, Geoffrey L. Herman, Mariana Silva, Maxwell Fowler, David H. Smith, Leo Porter 0001, Nico Ritschel, Craig B. Zilles, Matthew West 0001 |
SIGCSE (1) | 9 |
| 2026 | AI-Supported Grading and Rubric Refinement for Free Response QuestionsabstractManually grading free response questions remains a persistent challenge in education. While such questions offer valuable opportunities for student learning and critical thinking, their evaluation often requires substantial time and effort from instructors or teaching assistants. In addition to the grading workload, open-ended responses are susceptible to inconsistencies in scoring and may reflect unclear expectations, both of which can undermine the effectiveness and fairness of the assessment process. To address these challenges, we employed an AI-based grading system integrated in PrairieLearn to automatically evaluate student submissions to free response questions using a predefined set of rubric items. This approach not only streamlines the grading process but also enables direct comparison between AI-generated rubric applications and human judgments, providing insight into alignment and potential discrepancies. These discrepancies provided valuable insight, allowing us to iteratively revise and clarify the rubric items. Our experiences with using the AI grading system across several computing courses suggest that even experienced educators face difficulties articulating rubrics that are both specific and interpretable. We furthermore argue that more attention should be given to the iterative development and evaluation of rubrics. Chenyan Zhao, Maxwell Fowler, Yael Gertner, Seth Poulsen, Matthew West 0001, Mariana Silva |
SIGCSE (1) | 5 |
| 2025 | Frequent Testing vs. Second-chance Testing: An Exploration
Geoffrey L. Herman, Kajal Patel, Chinedu Emeka, Craig B. Zilles, Matthew West 0001 |
ICER (1) | 5 |
| 2025 | Measuring Test Anxiety of Two Computerized Exam ApproachesabstractComputerized exams have benefits for large enrollment courses and computer science classes, specifically. In this research paper, we compare student self-reported test anxiety between two modes of administering computerized exams: a computer-based testing facility (CBTF) and a bring-your-own-device (BYOD) setup. We conducted crossover design experiments in two computer science courses, measuring trait anxiety, as well as students' test anxiety and their test performance after each exam. Chinedu Emeka, Craig B. Zilles, Jim Sosnowski, Matthew West 0001, Geoffrey L. Herman, Mariana Silva |
SIGCSE (1) | 4 |
| 2025 | Measuring the Impact of Distractors on Student Learning Gains while Using Proof BlocksabstractBackground: Proof Blocks is a software tool that enables students to construct proofs by assembling prewritten lines and gives them automated feedback. Prior work on learning gains from Proof Blocks has focused on comparing learning gains from Proof Blocks against other learning activities such as writing proofs or reading. Seth Poulsen, Hongxuan Chen 0001, Yael Gertner, Benjamin Cosman, Matthew West 0001, Geoffrey L. Herman |
SIGCSE (1) | 5 |
| 2025 | Experiences with Computer-Based Testing (CBT)abstractDelivery of affordable, secure, and scalable assessments is an essential component of large university courses, whether online or in-person. The transition to Computer-Based Testing (CBT) has a transformational effect on pedagogy. Modern CBT systems provide almost unlimited flexibility in the types of questions they can support for manual grading and autograding. In this BoF, faculty interested in learning about various components of CBT and how to implement it at their institution are invited to ask their questions and learn from others who have already done this. To facilitate these discussions, in this BOF we will break into four smaller groups to discuss CBT pedagogy, building and sharing question banks, technical or logistical considerations, and building buy-in from all levels of the institution. Jim Sosnowski, Armando Fox, Dan Garcia 0001, Firas Moosvi, Mariana Silva, Matthew West 0001, Craig B. Zilles |
SIGCSE (2) | 6 |
| 2025 | Evaluating AI Models for Autograding Explain in Plain English Questions: Challenges and ConsiderationsabstractCode-reading ability has traditionally been under-emphasized in assessments as it is difficult to assess at scale. Prior research has shown that code-reading and code-writing are closely related skills; thus being able to assess and train code reading skills may be necessary for student learning. One way to assess code-reading ability is using Explain in Plain English (EiPE) questions, which ask students to describe what a piece of code does with natural language. Previous research deployed a binary (correct/incorrect) autograder using bigram models that performed comparably with human teaching assistants on student responses. With a dataset of 3,064 student responses from 17 EiPE questions, we investigated multiple autograders for EiPE questions. We evaluated methods as simple as logistic regression trained on bigram features, to more complicated Support Vector Machines (SVMs) trained on embeddings from Large Language Models (LLMs) to GPT-4. We found multiple useful autograders, most with accuracies in the \(86\!\!-\!\!88\%\) range, with different advantages. SVMs trained on LLM embeddings had the highest accuracy; few-shot chat completion with GPT-4 required minimal human effort; pipelines with multiple autograders for specific dimensions (what we call 3D autograders) can provide fine-grained feedback; and code generation with GPT-4 to leverage automatic code testing as a grading mechanism in exchange for slightly more lenient grading standards. While piloting these autograders in a non-major introductory Python course, students had largely similar views of all autograders, although they more often found the GPT-based grader and code-generation graders more helpful and liked the code-generation grader the most. Maxwell Fowler, Chinedu Emeka, Binglin Chen, David H. Smith IV, Matthew West 0001, Craig B. Zilles |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2024 | Learning from Integral Losses in Physics Informed Neural NetworksabstractThis work proposes a solution for the problem of training physics-informed networks under partial integro-differential equations. These equations require an infinite or a large number of neural evaluations to construct a single residual for training. As a result, accurate evaluation may be impractical, and we show that naive approximations at replacing these integrals with unbiased estimates lead to biased loss functions and solutions. To overcome this bias, we investigate three types of potential solutions: the deterministic sampling approaches, the double-sampling trick, and the delayed target method. We consider three classes of PDEs for benchmarking; one defining Poisson problems with singular charges and weak solutions of up to 10 dimensions, another involving weak solutions on electro-magnetic fields and a Maxwell equation, and a third one defining a Smoluchowski coagulation problem. Our numerical results confirm the existence of the aforementioned bias in practice and also show that our proposed delayed target approach can lead to accurate solutions with comparable quality to ones estimated with a large sample size integral. Our implementation is open-source and available at https://github.com/ehsansaleh/btspinn. Ehsan Saleh, Saba Ghaffari, Timothy Bretl, Luke N. Olson, Matthew West 0001 |
ICML | 5 |
| 2024 | A Comparison of Proctoring Regimens for Computer-Based Computer Science ExamsabstractIn this paper, we explore three different methods for administering computer-based tests at scale: (1) a dedicated Computer-Based Testing Center (CBTC), (2) Bring Your Own Device (BYOD) exams proctored in person in the classroom, and (3) BYOD exams proctored online via Zoom. We conducted two randomized crossover experiments to compare pairs of modalities against each other (CBTC vs BYOD-in-person and CBTC vs BYOD-online). We found that testing modality did not impact students' exam performance or students' preparation before exams. However, we observed that students preferred the modalities in which they had recently received the highest scores. Our results indicate that several different modalities can be effectively used to administer testing at scale for CS courses. Chinedu Emeka, Matthew West 0001, Craig B. Zilles, Mariana Silva |
ITiCSE (1) | 2 |
| 2024 | Plagiarism in the Age of Generative AI: Cheating Method Change and Learning Loss in an Intro to CS CourseabstractBackground: ChatGPT became widespread in early 2023 and enabled the broader public to use powerful generative AI, creating a new means for students to complete course assessments. Binglin Chen, Colleen M. Lewis, Matthew West 0001, Craig B. Zilles |
L@S | 3 |
| 2024 | Experiences With Computer-Based Testing (CBT)abstractAffordable, secure, and scalable assessment delivery is an essential component of large university courses, whether online or in-person. The switch to Computer-Based Testing (CBT) can have a surprising, and almost transformational effect on pedagogy. Modern CBT systems provide almost unlimited flexibility in the types of questions they can support, for both manual grading and autograding, and CBT has now been adopted at several universities and is under serious consideration at others. In this BoF, faculty interested in learning about CBT and how to implement it at their institution are invited to ask their questions. Faculty experienced with CBT are invited to share how CBT has changed their approach, pedagogy, and behavior and how to advocate for its adoption. Armando Fox, Dan Garcia 0001, Cinda Heeren, Firas Moosvi, Mariana Silva, Matthew West 0001, Craig B. Zilles |
SIGCSE (2) | 6 |
| 2024 | Comparing the Security of Three Proctoring Regimens for Bring-Your-Own-Device ExamsabstractWe compare the exam security of three proctoring regimens of Bring-Your-Own-Device, synchronous, computer-based exams in a computer science class: online un-proctored, online proctored via Zoom, and in-person proctored. We performed two randomized crossover experiments to compare these proctoring regimens. The first study measured the score advantage students receive while taking un-proctored online exams over Zoom-proctored online exams. The second study measured the score advantage of students taking Zoom-proctored online exams over in-person proctored exams. In both studies, students took six 50-minute exams using their own devices, which included two coding questions and 8--10 non-coding questions. We find that students score 2.3% higher on non-coding questions when taking exams in the un-proctored format compared to Zoom proctoring. No statistically significant advantage was found for the coding questions. While most of the non-coding questions had randomization such that students got different versions, for the few questions where all students received the same exact version, the score advantage escalated to 5.2%. From the second study, we find no statistically significant difference between students' performance on Zoom-proctored vs. in-person proctored exams. With this, we recommend educators incorporate some form of proctoring along with question randomization to mitigate cheating concerns in BYOD exams. Rishi Gulati, Matthew West 0001, Craig B. Zilles, Mariana Silva |
SIGCSE (1) | 2 |
| 2024 | Disentangling the Learning Gains from Reading a Book Chapter and Completing Proof Blocks ProblemsabstractBackground : Proof Blocks is a software tool that enables students to construct proofs by assembling prewritten lines and gives them automated feedback. Prior research has shown that students learn as much from an activity where they use Proof Blocks as where they write proofs. However, in both cases students first read a book chapter. Prior research was not able to differentiate between the learning gains achieved from reading versus proof practice. Purpose : This study aims to measure learning gains from reading a book chapter versus completing Proof Blocks. Methods : We conducted a randomized controlled trial with three experimental groups: one that only read a book chapter, one that only completed Proof Blocks, and one that did both. Findings : The group that completed only Proof Blocks had the smallest learning gains. The group that read the book chapter and completed the Proof Blocks activity performed marginally better than students who only read the book chapter, but it is not clear if the source of this improvement was the Proof Blocks or just exposure to more examples. Seth Poulsen, Yael Gertner, Hongxuan Chen 0001, Benjamin Cosman, Matthew West 0001, Geoffrey L. Herman |
SIGCSE (1) | 5 |
| 2023 | Efficient Feedback and Partial Credit Grading for Proof Blocks Problems
Seth Poulsen, Shubhang Kulkarni, Geoffrey L. Herman, Matthew West 0001 |
AIED | 4 |
| 2023 | Hierarchical Graph Neural Network with Cross-Attention for Cross-Device User Matching
Ali Taghibakhshi, Mingyuan Ma, Ashwath Aithal, Onur Yilmaz, Haggai Maron, Matthew West 0001 |
DaWaK | 6 |
| 2023 | MG-GNN: Multigrid Graph Neural Networks for Learning Multilevel Domain Decomposition MethodsabstractDomain decomposition methods (DDMs) are popular solvers for discretized systems of partial differential equations (PDEs), with one-level and multilevel variants. These solvers rely on several algorithmic and mathematical parameters, prescribing overlap, subdomain boundary conditions, and other properties of the DDM. While some work has been done on optimizing these parameters, it has mostly focused on the one-level setting or special cases such as structured-grid discretizations with regular subdomain construction. In this paper, we propose multigrid graph neural networks (MG-GNN), a novel GNN architecture for learning optimized parameters in two-level DDMs. We train MG-GNN using a new unsupervised loss function, enabling effective training on small problems that yields robust performance on unstructured grids that are orders of magnitude larger than those in the training set. We show that MG-GNN outperforms popular hierarchical graph network architectures for this optimization and that our proposed loss function is critical to achieving this improved performance. Ali Taghibakhshi, Nicolas Nytko, Tareq Uz Zaman, Scott P. MacLachlan, Luke N. Olson, Matthew West 0001 |
ICML | 6 |
| 2023 | A's for All (As Time and Interest Allow)abstract"A's for All (as time and interest allow)" is a position that says it is increasingly possible to aim for a world in which students can achieve any grade (level of mastery) that they are willing to work for, even if some students take longer than others or require more practice to get there. Achieving this goal would have profound effects on fairness, equity, and participation in computing, to say nothing of student learning outcomes. We describe what this goal would entail, why it is worth pursuing, what the mechanism and policy requirements are for making progress, and why now is a good time to do it. We give specific and actionable recommendations, many based on our own experience so far, that our colleagues who are excited about the approach can put into immediate practice, and address a number of concerns and objections that our proposal may raise. Importantly, our proposed approach is not all-or-nothing, but all-or-something: there are many things instructors can do within existing policy frameworks and course constraints to move their course experience in this direction. Dan Garcia 0001, Armando Fox, Solomon Russell, Edwin Ambrosio, Neal Terrell, Mariana Silva, Matthew West 0001, Craig B. Zilles, Fuzail Shakir |
SIGCSE (1) | 7 |
| 2023 | Actually Achieving "A's for All" (As Time and Interest Allow)abstractIn recent years, a diverse body of research in computing education has discussed new pedagogies and curriculum changes to improve learning and students' experiences. Topics such as growth mindset, mastery learning, grading for equity, and specifications grading are important steps towards the Holy Grail: "A's for All" (as time and interest allow). In this new teaching approach, the "A" line does not move, but instead every student is given an opportunity to achieve proficiency and earn it, as long as they are willing to put in the time and effort it takes. In other words, students can all achieve the same learning outcomes at a different pace, instead of the traditional approach where students achieve different learning outcomes in a fixed amount of time. This workshop will provide educators and administrators with tools to implement the "A's for All" (as time and interest allow) approach in their courses and institutions. It will take attendees through elements of advocacy, hands-on randomized question generator design and implementation using a computer-based assessment system, best practices, and course policies to reduce friction. Dan Garcia 0001, Connor McMahon, Yuan Garcia, Craig B. Zilles, Matthew West 0001, Mariana Silva, Solomon Russell, Edwin Ambrosio, Neal Terrell |
SIGCSE (2) | 5 |
| 2023 | Measuring the Impact of a Computational Linear Algebra Course on Students' Exam Performance in a Subsequent Numerical Methods CourseabstractA new computational linear algebra course was developed and offered at a large public university in the Midwest. This new course traded off some of the lecture time in the pre-existing traditional linear algebra course for applied computational materials taught in a flipped-classroom lab setting. We compare exam performance in a subsequent numerical methods course from students having taken either the new computational or traditional course, while controlling for student performance in prerequisite computer science and mathematics courses. We find that for students with less mathematics background (i.e., those who needed to take Calculus 2 at the university), taking the new computational linear algebra course has significant positive impact on their average exam performance in the subsequent course. The performance of students with more initial mathematics background (i.e., those who already had credit for Calculus 2) is not significantly affected by the computational vs. traditional course backgrounds. Hongxuan Chen 0001, Matthew West 0001, Sascha Hilgenfeldt, Mariana Silva |
SIGCSE (1) | 2 |
| 2023 | Efficiency of Learning from Proof Blocks Versus Writing ProofsabstractProof Blocks is a software tool that provides students with a scaffolded proof-writing experience, allowing them to drag and drop prewritten proof lines into the correct order instead of starting from scratch. In this paper we describe a randomized controlled trial designed to measure the learning gains of using Proof Blocks for students learning proof by induction. The study participants were 332 students recruited after completing the first month of their discrete mathematics course. Students in the study took a pretest and read lecture notes on proof by induction, completed a brief (less than 1 hour) learning activity, and then returned one week later to complete the posttest. Depending on the experimental condition that each student was assigned to, they either completed only Proof Blocks problems, completed some Proof Blocks problems and some written proofs, or completed only written proofs for their learning activity. We find that students in the early phases of learning about proof by induction are able to learn just as much from reading lecture notes and using Proof Blocks as by reading lecture notes and writing proofs from scratch, but in far less time on task. This finding complements previous findings that Proof Blocks are useful exam questions and are viewed positively by students. Seth Poulsen, Yael Gertner, Benjamin Cosman, Matthew West 0001, Geoffrey L. Herman |
SIGCSE (1) | 4 |
| 2023 | Investigating the Effects of Testing Frequency on Programming Performance and Students' BehaviorabstractWe conducted an across-semester quasi-experimental study that compared students' outcomes under frequent and infrequent testing regimens in an introductory computer science course. Students in the frequent testing (4 quizzes and 4 exams) semester outperformed the infrequent testing (1 midterm and 1 final exam) semester by 9.1 to 13.5 percentage points on code writing questions. David H. Smith IV, Chinedu Emeka, Maxwell Fowler, Matthew West 0001, Craig B. Zilles |
SIGCSE (1) | 4 |
| 2022 | Proof Blocks: Autogradable Scaffolding Activities for Learning to Write ProofsabstractIn this software tool paper we present Proof Blocks, a tool which enables students to construct mathematical proofs by dragging and dropping prewritten proof lines into the correct order. We present both implementation details of the tool, as well as a rich reflection on our experiences using the tool in courses with hundreds of students. Proof Blocks problems can be graded completely automatically, enabling students to receive rapid feedback. When writing a problem, the instructor specifies the dependency graph of the lines of the proof, so that any correct arrangement of the lines can receive full credit. This innovation can improve assessment tools by increasing the types of questions we can ask students about proofs, and can give greater access to proof knowledge by increasing the amount that students can learn on their own with the help of a computer. Seth Poulsen, Mahesh Viswanathan 0001, Geoffrey L. Herman, Matthew West 0001 |
ITiCSE (1) | 4 |
| 2022 | Achieving "A's for All (as Time and Interest Allow)"abstractThe SIGCSE-MEMBERS mailing list of the ACM Special Interest Group in Computer Science Education is the main forum for educators worldwide to discuss computing education research, pedagogy, and curriculum. In early 2022 it was abuzz with several connected movements: growth mindset, proficiency (aka mastery) learning, grading for equity, and specifications grading. Each of these is an important step toward the Holy Grail: A's for All (as time and interest allow); the "A" line doesn't move, but every student should be given an opportunity to achieve proficiency and earn it, as long as they are willing to put in the time and effort it might take. The mantra is not "fixed time, variable learning", but "fixed learning, variable time". The goal of this full-day workshop is to provide educators and administrators with tools to achieve it in their courses and institutions. Dan Garcia 0001, Connor McMahon, Yuan Garcia, Matthew West 0001, Craig B. Zilles |
L@S | 4 |
| 2022 | Software Support for "A's for All"abstractThe SIGCSE-MEMBERS mailing list of the ACM Special Interest Group in Computer Science Education is the main forum for educators worldwide to discuss computing education research, pedagogy, and curriculum. In early 2022 it was abuzz with several connected movements: growth mindset, proficiency (aka mastery) learning, grading for equity, and specifications grading. Each of these is an important step toward the Holy Grail: A's for All (as time and interest allow); the "A" line doesn't move, but every student should be given an opportunity to achieve proficiency and earn it, as long as they are willing to put in the time and effort it might take. The mantra is not "fixed time, variable learning", but "fixed learning, variable time". Dan Garcia 0001, Connor McMahon, Yuan Garcia, Matthew West 0001, Craig B. Zilles |
L@S | 4 |
| 2022 | Truly Deterministic Policy OptimizationabstractIn this paper, we present a policy gradient method that avoids exploratory noise injection and performs policy search over the deterministic landscape, with the goal of improving learning with long horizons and non-local rewards. By avoiding noise injection all sources of estimation variance can be eliminated in systems with deterministic dynamics (up to the initial state distribution). Since deterministic policy regularization is impossible using traditional non-metric measures such as the KL divergence, we derive a Wasserstein-based quadratic model for our purposes. We state conditions on the system model under which it is possible to establish a monotonic policy improvement guarantee, propose a surrogate function for policy gradient estimation, and show that it is possible to compute exact advantage estimates if both the state transition model and the policy are deterministic. Finally, we describe two novel robotic control environments---one with non-local rewards in the frequency domain and the other with a long horizon (8000 time-steps)---for which our policy gradient method (TDPO) significantly outperforms existing methods (PPO, TRPO, DDPG, and TD3). Our implementation with all the experimental settings and a video of the physical hardware test is available at https://github.com/ehsansaleh/tdpo . Ehsan Saleh, Saba Ghaffari, Timothy Bretl, Matthew West 0001 |
NeurIPS | 4 |
| 2022 | Learning Interface Conditions in Domain Decomposition SolversabstractDomain decomposition methods are widely used and effective in the approximation of solutions to partial differential equations. Yet the \textit{optimal} construction of these methods requires tedious analysis and is often available only in simplified, structured-grid settings, limiting their use for more complex problems. In this work, we generalize optimized Schwarz domain decomposition methods to unstructured-grid problems, using Graph Convolutional Neural Networks (GCNNs) and unsupervised learning to learn optimal modifications at subdomain interfaces. A key ingredient in our approach is an improved loss function, enabling effective training on relatively small problems, but robust performance on arbitrarily large problems, with computational cost linear in problem size. The performance of the learned linear solvers is compared with both classical and optimized domain decomposition algorithms, for both structured- and unstructured-grid problems. Ali Taghibakhshi, Nicolas Nytko, Tareq Uz Zaman, Scott P. MacLachlan, Luke N. Olson, Matthew West 0001 |
NeurIPS | 6 |
| 2022 | Peer-grading "Explain in Plain English": A Bayesian Calibration Method for Categorical Answersabstract"Explain in plain English'' (EipE) questions have been proposed as an important activity and assessment for studying novice programmers' grasp of programming knowledge and their ability to communicate their understanding. However, EipE questions aren't widely used in introductory programming courses in part because of the large grading effort required. In this paper, we present our experience of using peer grading for EipE questions in a large-enrollment introductory programming course, where students were asked to categorize other students' responses. We developed a novel Bayesian algorithm for performing calibrated peer grading on categorical data, and we used a heuristic grade assignment method based on the Bayesian estimates. The peer-grading exercises served both as a way to coach students on what is expected from EipE questions and as a way to alleviate the grading load for the course staff. Based on four rounds of peer-grading activities, we found that students are generally capable of categorizing responses to EiPE questions and that our proposed Bayesian method is more robust than unweighted voting. Binglin Chen, Matthew West 0001, Craig B. Zilles |
SIGCSE (1) | 2 |
| 2022 | Are We Fair?: Quantifying Score Impacts of Computer Science Exams with Randomized Question PoolsabstractWith the increase of large enrollment courses and the growing need to offer online instruction, computer-based exams randomly generated from question pools have a clear benefit for computing courses. Such exams can be used at scale, scheduled asynchronously and/or online, and use versioning to make attempts at cheating less profitable. Despite these benefits, we want to ensure that the technique is not unfair to students, particularly when it comes to equivalent difficulty across exam versions. Maxwell Fowler, David H. Smith IV, Chinedu Emeka, Matthew West 0001, Craig B. Zilles |
SIGCSE (1) | 4 |
| 2021 | Students' Perceptions and Behavior Related to Second-Chance TestingabstractThis full research paper explores students' attitudes toward second-chance testing and how second-chance testing influences students' behavior. Second-chance testing refers to giving students the opportunity to take a second instance of each exam for some sort of grade replacement. Previous work has demonstrated that second-chance testing can lead to improved student outcomes in courses, but how to best structure second-chance testing to maximize its benefits remains an open question. We complement previous work by interviewing a diverse group of 23 students that have taken courses that use second-chance testing. From the interviews, we sought to gain insight into students' views and use of second-chance testing. We found that second-chance testing was almost universally viewed positively by the students and was frequently cited as helping to reduce test takers' anxiety and boost their confidence. Overall, we find that the majority of students prepare for second-chance exams in desirable ways, but we also note ways in which second-chance testing can potentially lead to undesirable behaviors including procrastination, overreliance on memorization, and attempts to game the system. We identified emergent themes pertaining to various facets of second-chance test-taking, including: 1) concerns about the time commitment required for second-chance exams; 2) a belief that second-chance exams promoted fairness; and 3) how second-chance testing incentivized learning. This paper will provide instructors and other stakeholders with detailed insights into students' behavior regarding second-chance testing, enabling instructors to develop better policies and avoid unintended consequences. Chinedu Emeka, Timothy Bretl, Geoffrey L. Herman, Matthew West 0001, Craig B. Zilles |
FIE | 4 |
| 2021 | Evaluating Proof Blocks Problems as Exam QuestionsabstractProof Blocks is a novel software tool which enables students to write mathematical proofs by dragging and dropping prewritten lines into the correct order, rather than writing a proof completely from scratch. We used Proof Blocks problems as exam questions for a discrete mathematics course with hundreds of students, allowing us to collect thousands of student responses to Proof Blocks problems. Using this data, we provide statistical evidence that Proof Blocks are easier than written proofs, which are typically very difficult. We also show that Proof Blocks problems provide about as much information about student knowledge as written proofs. Survey results show that students believe that the Proof Blocks user interface is easy to use, and that the questions accurately represent their ability to write proofs. Seth Poulsen, Mahesh Viswanathan 0001, Geoffrey L. Herman, Matthew West 0001 |
ICER | 4 |
| 2021 | Optimization-Based Algebraic Multigrid Coarsening Using Reinforcement LearningabstractLarge sparse linear systems of equations are ubiquitous in science and engineering, such as those arising from discretizations of partial differential equations. Algebraic multigrid (AMG) methods are one of the most common methods of solving such linear systems, with an extensive body of underlying mathematical theory. A system of linear equations defines a graph on the set of unknowns and each level of a multigrid solver requires the selection of an appropriate coarse graph along with restriction and interpolation operators that map to and from the coarse representation. The efficiency of the multigrid solver depends critically on this selection and many selection methods have been developed over the years. Recently, it has been demonstrated that it is possible to directly learn the AMG interpolation and restriction operators, given a coarse graph selection. In this paper, we consider the complementary problem of learning to coarsen graphs for a multigrid solver, a necessary step in developing fully learnable AMG methods. We propose a method using a reinforcement learning (RL) agent based on graph neural networks (GNNs), which can learn to perform graph coarsening on small planar training graphs and then be applied to unstructured large planar graphs, assuming bounded node degree. We demonstrate that this method can produce better coarse graphs than existing algorithms, even as the graph size increases and other properties of the graph are varied. We also propose an efficient inference procedure for performing graph coarsening that results in linear time complexity in graph size. Ali Taghibakhshi, Scott P. MacLachlan, Luke N. Olson, Matthew West 0001 |
NeurIPS | 4 |
| 2021 | AutogradingabstractPrevious research suggests that "Explain in Plain English" (EiPE) code reading activities could play an important role in the development of novice programmers, but EiPE questions aren't heavily used in introductory programming courses because they (traditionally) required manual grading. We present what we believe to be the first automatic grader for EiPE questions and its deployment in a large-enrollment introductory programming course. Based on a set of questions deployed on a computer-based exam, we find that our implementation has an accuracy of 87-89%, which is similar in performance to course teaching assistants trained to perform this task and compares favorably to automatic short answer grading algorithms developed for other domains. In addition, we briefly characterize the kinds of answers that the current autograder fails to score correctly and the kinds of errors made by students. Maxwell Fowler, Binglin Chen, Sushmita Azad, Matthew West 0001, Craig B. Zilles |
SIGCSE | 4 |
| 2021 | Verifying Stochastic Hybrid Systems with Temporal Logic Specifications via Model ReductionabstractWe present a scalable methodology to verify stochastic hybrid systems for inequality linear temporal logic (iLTL) or inequality metric interval temporal logic (iMITL). Using the Mori–Zwanzig reduction method, we construct a finite-state Markov chain reduction of a given stochastic hybrid system and prove that this reduced Markov chain is approximately equivalent to the original system in a distributional sense. Approximate equivalence of the stochastic hybrid system and its Markov chain reduction means that analyzing the Markov chain with respect to a suitably strengthened property allows us to conclude whether the original stochastic hybrid system meets its temporal logic specifications. Based on this, we propose the first statistical model checking algorithms to verify stochastic hybrid systems against correctness properties, expressed in iLTL or iMITL. The scalability of the proposed algorithms is demonstrated by a case study. Yu Wang 0044, Nima Roohi, Matthew West 0001, Mahesh Viswanathan 0001, Geir E. Dullerud |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Strategies for Deploying Unreliable AI Graders in High-Transparency High-Stakes Exams
Sushmita Azad, Binglin Chen, Maxwell Fowler, Matthew West 0001, Craig B. Zilles |
AIED (1) | 4 |
| 2020 | STMC: Statistical Model Checker with Stratified and Antithetic Samplingabstractis a statistical model checker that uses antithetic and stratified sampling techniques to reduce the number of samples and, hence, the amount of time required before making a decision. The tool is capable of statistically verifying any black-box probabilistic system that can simulate, against probabilistic bounds on any property that can evaluate over individual executions of the system. We have evaluated our tool on many examples and compared it with both symbolic and statistical algorithms. When the number of strata is large, our algorithms reduced the number of samples more than 3 times on average. Furthermore, being a statistical model checker makes able to verify models that are well beyond the reach of current symbolic model checkers. On large systems (up to $$10^{14}$$ states) was able to check 100% of benchmark systems, compared to existing symbolic methods in , which only succeeded on 13% of systems. The tool, installation instructions, benchmarks, and scripts for running the benchmarks are all available online as open source. Nima Roohi, Yu Wang 0044, Matthew West 0001, Geir E. Dullerud, Mahesh Viswanathan 0001 |
CAV (2) | 3 |
| 2020 | Comparison of Grade Replacement and Weighted Averages for Second-Chance ExamsabstractWe explore how course policies affect students' studying and learning when a second-chance exam is offered. High-stakes, one-off exams remain a de facto standard for assessing student knowledge in STEM, despite compelling evidence that other assessment paradigms such as mastery learning can improve student learning. Unfortunately, mastery learning can be costly to implement. We explore the use of optional second-chance testing to sustainably reap the benefits of mastery-based learning at scale. Prior work has shown that course policies affect students' studying and learning but have not compared these effects within the same course context. We conducted a quasi-experimental study in a single course to compare the effect of two grading policies for second-chance exams and the effect of increasing the size of the range of dates for students taking asynchronous exams. The first grading policy, called 90-cap, allowed students to optionally take a second-chance exam that would fully replace their score on a first-chance exam except the second-chance exam would be capped at 90% credit. The second grading policy, called 90-10, combined students' first- and second-chance exam scores as a weighted average (90% max score + 10% min score). The 90-10 policy significantly increased the likelihood that marginally competent students would take the second-chance exam. Further, our data suggests that students learned more under the 90-10 policy, providing improved student learning outcomes at no cost to the instructor. Most students took exams on the last day an exam was available, regardless of how many days the exam was available. Geoffrey L. Herman, Zhouxiang Cai, Timothy Bretl, Craig B. Zilles, Matthew West 0001 |
ICER | 5 |
| 2020 | Learning to Cheat: Quantifying Changes in Score Advantage of Unproctored Assessments Over TimeabstractProctoring educational assessments (e.g., quizzes and exams) has a cost, be it in faculty (and/or course staff) time or in money to pay for proctoring services. Previous estimates of the utility of proctoring (generally by estimating the score advantage of taking an exam without proctoring) vary widely and have mostly been implemented using an across subjects experimental designs and sometimes with low statistical power. Binglin Chen, Sushmita Azad, Maxwell Fowler, Matthew West 0001, Craig B. Zilles |
L@S | 4 |
| 2020 | A Quantitative Analysis of When Students Choose to Grade Questions on Computerized Exams with Multiple AttemptsabstractIn this paper, we study a computerized exam system that allows students to attempt the same question multiple times. This system permits students either to receive feedback on their submitted answer immediately or to defer the feedback and grade questions in bulk. An analysis of student behavior in three courses across two semesters found similar student behaviors across courses and student groups. We found that only a small minority of students used the deferred feedback option. A clustering analysis that considered both when students chose to receive feedback and either to immediately retry incorrect problems or to attempt other unfinished problems identified four main student strategies. These strategies were correlated to statistically significant differences in exam scores, but it was not clear if some strategies improved outcomes or if stronger students tended to prefer certain strategies. Ashank Verma, Timothy Bretl, Matthew West 0001, Craig B. Zilles |
L@S | 3 |
| 2020 | A Validated Scoring Rubric for Explain-in-Plain-English QuestionsabstractPrevious research has identified the ability to read code and understand its high-level purpose as an important developmental skill that is harder to do (for a given piece of code) than executing code in one's head for a given input ("code tracing"), but easier to do than writing the code. Prior work involving code reading ("Explain in plain English") problems, have used a scoring rubric inspired by the SOLO taxonomy, but we found it difficult to employ because it didn't adequately handle the three dimensions of answer quality: correctness, level of abstraction, and ambiguity. In this paper, we describe a 7-point rubric that we developed for scoring student responses to "Explain in plain English'' questions, and we validate this rubric through four means. First, we find that the scale can be reliably applied with with a median Krippendorff's alpha (inter-rater reliability) of 0.775. Second, we report on an experiment to assess the validity of our scale. Third, we find that a survey consisting of 12 code reading questions had a high internal consistency (Cronbach's alpha = 0.954). Last, we find that our scores for code reading questions in a large enrollment (N = 452) data structures course are correlated (Pearson's R = 0.555) to code writing performance to a similar degree as found in previous work. Binglin Chen, Sushmita Azad, Rajarshi Haldar, Matthew West 0001, Craig B. Zilles |
SIGCSE | 4 |
| 2020 | Measuring the Score Advantage on Asynchronous Exams in an Undergraduate CS CourseabstractThis paper presents the results of a controlled crossover experiment designed to measure the score advantage that students have when taking exams asynchronously (i.e., the students can select a time to take the exam in a multi-day window) compared to synchronous exams (i.e., all students take the exam at the same time). The study was performed in an upper-division undergraduate computer science course with 321 students. Stratified sampling was used to randomly assign the students to two groups that alternated between the two treatments (synchronous versus asynchronous exams) across a series of four exams during the semester. These non-programming exams consisted of a mix of multiple choice, checkbox, and numeric input questions. For some questions, the parameters were randomized so that students received different versions of the question and some questions were identical for all students. In our results, students taking the exams asynchronously had scores that were on average only 3% higher (0.2 of a standard deviation). Furthermore, we found that the score advantage was decreased by the use of randomized questions, and it did not significantly differ based on the type of question. Thus, our results suggest that asynchronous exams can be a compelling alternative to synchronous exams. Mariana Silva, Matthew West 0001, Craig B. Zilles |
SIGCSE | 2 |
| 2019 | Effect of Discrete and Continuous Parameter Variation on Difficulty in Automatic Item Generation
Binglin Chen, Craig B. Zilles, Matthew West 0001, Timothy Bretl |
AIED (1) | 3 |
| 2019 | Every University Should Have a Computer-Based Testing FacilityabstractFor the past five years we have been operating a Computer-Based Testing Facility (CBTF) as the primary means of summative assessment in large-enrollment STEM-oriented classes. In each of the last three semesters, it has proctored over 50,000 exams for over 6,000 unique students in 25–30 classes. Our CBTF has simultaneously improved the quality of assessment, allowed the testing of computational skills, and reduced the recurring burden of performing assessment in a broad collection of STEM-oriented classes, but it does require an up-front investment to develop the digital exam content. We have found our CBTF to be secure, cost-effective, and well liked by our faculty, who choose to use it semester after semester. We believe that there are many institutions that would similarly benefit from having a Computer-Based Testing Facility. Craig B. Zilles, Matthew West 0001, Geoffrey L. Herman, Timothy Bretl |
CSEDU (1) | 2 |
| 2019 | Predicting the difficulty of automatic item generators on exams from their difficulty on homeworksabstractTo design good assessments, it is useful to have an estimate of the difficulty of a novel exam question before running an exam. In this paper, we study a collection of a few hundred automatic item generators (short computer programs that generate a variety of unique item instances) and show that their exam difficulty can be roughly predicted from student performance on the same generator during pre-exam practice. Specifically, we show that the rate that students correctly respond to a generator on an exam is on average within 5% of the correct rate for those students on their last practice attempt. This study is conducted with data from introductory undergraduate Computer Science and Mechanical Engineering courses. Binglin Chen, Matthew West 0001, Craig B. Zilles |
L@S | 2 |
| 2019 | Statistical verification of PCTL using antithetic and stratified samples
Yu Wang 0044, Nima Roohi, Matthew West 0001, Mahesh Viswanathan 0001, Geir E. Dullerud |
Formal Methods Syst. Des. | 3 |
| 2018 | Towards a Model-Free Estimate of the Limits to Student Modeling Accuracy
Binglin Chen, Matthew West 0001, Craig B. Zilles |
EDM | 2 |
| 2018 | Making Testing Less Trying: Lessons Learned from Operating a Computer-Based Testing FacilityabstractThis Innovative Practice Full Paper describes lessons learned from and operational details of a full-scale Computer-Based Testing Facility (CBTF) over a period of almost 4 years. The CBTF has grown into a key resource for enabling the graceful scaling of many of the largest classes in our College of Engineering. In Fall 2017, the CBTF served 21 courses from seven different departments and over 6,000 unique students. Over 52,000 exams were delivered, including 3,500 final exams.This paper discusses five main aspects of our CBTF. First, we present the basic operation of the CBTF. Second, we discuss the precautions we take to maintain a secure exam environment. Third, we discuss how we support students that require testing accommodations like extra time and/or a distraction-reduced environment. Fourth, we discuss how we organize our policies to handle exceptional circumstances with minimal intervention by faculty. Finally, we discuss the cost of operating the CBTF and how it compares to traditional exams and online services. Craig B. Zilles, Matthew West 0001, David Mussulman, Timothy Bretl |
FIE | 2 |
| 2018 | How much randomization is needed to deter collaborative cheating on asynchronous exams?abstractThis paper investigates randomization on asynchronous exams as a defense against collaborative cheating. Asynchronous exams are those for which students take the exam at different times, potentially across a multi-day exam period. Collaborative cheating occurs when one student (the information producer) takes the exam early and passes information about the exam to other students (the information consumers) that are taking the exam later. Using a dataset of computerized exam and homework problems in a single course with 425 students, we identified 5.5% of students (on average) as information consumers by their disproportionate studying of problems that were on the exam. These information consumers ("cheaters") had a significant advantage (13 percentage points on average) when every student was given the same exam problem (even when the parameters are randomized for each student), but that advantage dropped to almost negligible levels (2--3 percentage points) when students were given a random problem from a pool of two or four problems. We conclude that randomization with pools of four (or even three) problems, which also contain randomized parameters, is an effective mitigation for collaborative cheating. Our analysis suggests that this mitigation is in part explained by cheating students having less complete information about larger pools. Binglin Chen, Matthew West 0001, Craig B. Zilles |
L@S | 2 |
| 2018 | Using a Computer-based Testing Facility to Improve Student Learning in a Programming Languages and Compilers CourseabstractWhile most efforts to improve students' learning in computer science education have focused on designing new pedagogies or tools, comparatively little research has focused on redesigning examinations to improve students' learning. Cognitive science research, however, has robustly demonstrated that getting students to practice using their knowledge in testing environments can significantly improve learning through a phenomenon known as the testing effect. The testing effect has been shown to improve learning more than rehearsal strategies such as re-reading a textbook or re-watching lectures. In this paper, we present a quasi-experimental study to examine the effect of using frequent, automated examinations in an advanced computer science course, "Programming Languages and Compilers" (CS 421). In Fall 2014, students were given traditional paper-based exams, but in Fall 2015 a computer-based testing facility enabled the course to offer more frequent examinations while other aspects of the course were held constant. A comparison of 292 student scores across the two semesters revealed a significant change in the distribution of students' grades with fewer students failing the final examination, and proportionately more students now earning grades of B and C instead. This data suggests that focusing on redesigning the nature of examinations may indeed be a relatively untapped opportunity to improve students' learning. Terence Nip, Elsa L. Gunter, Geoffrey L. Herman, Jason Morphew, Matthew West 0001 |
SIGCSE | 5 |
| 2018 | An Improved Grade Point Average, With Applications to CS Undergraduate Education AnalyticsabstractWe present a methodological improvement for calculating Grade Point Averages (GPAs). Heterogeneity in grading between courses systematically biases observed GPAs for individual students: the GPA observed depends on course selection. We show how a logistic model can account for course selection by simulating how every student in a sample would perform if they took all available courses, giving a new “modeled GPA.” We then use 10 years of grade data from a large university to demonstrate that this modeled GPA is a more accurate predictor of student performance in individual courses than the observed GPA. Using Computer Science (CS) as an example learning analytics application, it is found that required CS courses give significantly lower grades than average courses. This depresses the recorded GPAs of CS majors: modeled GPAs are 0.25 points higher than those that are observed. The modeled GPA also correlates much more closely with standardized test scores than the observed GPA: the correlation with Math ACT is 0.37 for the modeled GPA and is 0.20 for the observed GPA. This implies that standardized test scores are much better predictors of student performance than might otherwise be assumed. Jonathan H. Tomkin, Matthew West 0001, Geoffrey L. Herman |
ACM Trans. Comput. Educ. | 2 |
| 2017 | Statistical Verification of the Toyota Powertrain Control Verification BenchmarkabstractThe Toyota Powertrain Control Verification Benchmark has been recently proposed as challenge problems that capture features of realistic automotive designs. In this paper we statistically verify the most complicated of the powertrain control models proposed, that includes features like delayed differential and difference equations, look-up tables, and highly non-linear dynamics, by simulating the C++ code generated from the SimulinkTM model of the design. Our results show that for at least 98% of the possible initial operating conditions the desired properties hold. These are the first verification results for this model, statistical or otherwise. Nima Roohi, Yu Wang 0044, Matthew West 0001, Geir E. Dullerud, Mahesh Viswanathan 0001 |
HSCC | 3 |
| 2017 | Do Performance Trends Suggest Wide-spread Collaborative Cheating on Asynchronous Exams?abstractUsing a data set from 29,492 asynchronous exams in an on-campus proctored computer-based testing facility (CBTF), we observed correlations between when a student chooses to take their exam within the exam period and their score on the exam. Somewhat surprisingly, instead of increasing throughout the exam period, which might be indicative of widespread collaborative cheating, we find that exam scores decrease throughout the exam period. While this could be attributed to weaker students putting off exams, this effect holds even when accounting for student ability as measured by a synchronous exam taken during the same semester. This suggests that precautions can be taken by a CBTF to maintain cheating at a low level (e.g., the level of proctored synchronous exams), in spite of the fact that students are taking their exams over a multi-day period. Binglin Chen, Matthew West 0001, Craig B. Zilles |
L@S | 2 |
| 2016 | Studying faculty Communities of Practice through social network analysisabstractCreating systemic change in undergraduate engineering and STEM education is difficult to achieve and just as difficult to study. It has been proposed that organizational learning and change theories can be coupled with social network analysis to achieve both of these goals. In this paper, we describe an institutional change effort designed around principles from Communities of Practice. We then present the design of a social analysis network study that we are executing to study and analyze whether this change effort has been successful in achieving its goals. We present some preliminary data to demonstrate the promise of this approach for executing and studying institutional change in engineering education and STEM education more broadly. Shufeng Ma, Matthew West 0001, Geoffrey L. Herman, Jonathan H. Tomkin, Jose Mestre |
FIE | 2 |
| 2016 | A methodological refinement for studying the STEM grade-point penaltyabstractWe present a study that explores the grade-point average (GPA) penalties that students face when taking introductory STEM courses. Previous work has found that there is a large and significant grade point penalty for women and minorities in most STEM classes (that is, these students perform worse in these classes than their overall GPA would suggest). We recreated this work using a new, large data set (63,012 students over 10 years) of student performance, and found that the initial results held when using the original approach. We argue that there are methodological shortcomings to the original approach, however, as there is no attempt to control for individual student program difficulty (STEM majors and non-STEM majors share some classes, but have very different overall suites of courses that determine their overall GPA). As the female/male and racial ratios vary across majors it is therefore likely that a division by gender is not comparing equivalent sample populations. By controlling for student test scores or major most of the penalty is removed. The initial findings of large GPA penalties in STEM courses appears to be an example of “Simpson's Paradox”. Jonathan H. Tomkin, Matthew West 0001, Geoffrey L. Herman |
FIE | 2 |
| 2016 | Modeling Student Scheduling Preferences in a Computer-Based Testing FacilityabstractWhen undergraduate students are allowed to choose a time slot in which to take an exam from a large number of options (e.g., 40), the students exhibit strong preferences among the times. We found that students can be effectively modelled using constrained discrete choice theory to quantify these preferences from their observed behavior. The resulting models are suitable for load balancing when scheduling multiple concurrent exams and for capacity planning given a set schedule. Matthew West 0001, Craig B. Zilles |
L@S | 1 |
| 2015 | Statistical verification of dynamical systems using set oriented methodsabstractModeling, analyzing and verifying real physical systems has long been a challenging task since the state space of the systems is usually infinite and the dynamics of the systems is generally nonlinear and stochastic. In this work, we employ an extension of linear temporal logic (LTL) to describe the behavior of discrete-time nonlinear stochastic systems; this extension is so-called linear inequality LTL (iLTL) which allows for atomic propositions that are linear inequalities on state spaces. To statistically verify iLTL formulas on the systems, we first reformulate discrete-time nonlinear stochastic dynamical systems into Markov processes on their continuous state spaces and then reduce them to discrete-time Markov chains (DTMC) using set-oriented methods. Furthermore, a statistical verification algorithm is proposed to verify iLTL formulas on the reduced systems. The correctness of this statistical verification algorithm is checked both by theoretical analysis and the simulation of a fluid problem. We will show in the successive work that the framework extends to hybrid systems, which is a significant motivation for the approach taken. Yu Wang 0044, Nima Roohi, Matthew West 0001, Mahesh Viswanathan 0001, Geir E. Dullerud |
HSCC | 3 |