VLDB 2026 Research / reviewers in the wild / expert
Marshall An
dblp:243/3793
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0009-0005-5165-640XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting the Hint Button: Consistent Negative Associations Between Unproductive Hint Use and Learning Outcomes in Intelligent Tutoring Systems
Marshall An, Mahboobeh Mehrvarz, John C. Stamper, Bruce M. McLaren |
LAK | 1 |
| 2026 | Deriving Instructional Insights from Human-LLM Co-Evaluation of Student Collaboration in Data-Centric ProgrammingabstractThis quasi-experimental study integrates a large language model (LLM) with expert qualitative analysis to examine how instructional design variations in computer-supported collaborative learning (CSCL) shape collaboration in data-centric programming. We collected 73 team transcripts from two contrasting CSCL designs deployed across five course offerings: a closed-ended variant with prescribed solution paths and auto-graded milestones, and an open-ended variant supporting exploratory tasks with multiple valid paths. LLM annotation revealed statistically significant differences in knowledge co-construction patterns: the open-ended design yielded a higher proportion of utterances focused on developing a shared understanding of problems and solutions. Guided by these quantitative results, human experts conducted qualitative coding that confirmed and enriched these findings, showing how open-ended tasks fostered elaborative solution negotiation while closed-ended structures promoted non-elaborative exchanges. Our contributions are: (1) instructional insights for data science education, demonstrating how open-ended CSCL designs better support collaborative sense-making essential for real-world data science projects; and (2) a documented workflow for human-LLM co-evaluation, providing the methodological detail necessary for others to replicate our process and apply it to future studies. Marshall An, Christine Kwon, Jihyeon Hur, Dongho Lee, Vincent Huai, Barry Zheng, Matthew Yu, Joana Liu, Jenny Pugh, Gahgene Gweon, John C. Stamper |
SIGCSE (1) | 1 |
| 2026 | Partnering with Community College Faculty to Co-Design Intelligent Tutoring Systems for Cybersecurity Workforce TrainingabstractThis experience report describes a partnership between community college faculty and learning scientists to co-design Intelligent Tutoring Systems (ITSs) addressing challenges in cybersecurity workforce training. Our co-design approach combined collaborative reflection on student difficulties from prior course offerings with systematic curricular analysis to identify high-impact intervention points. We targeted two challenge areas: strengthening students' ability to contrast key cybersecurity taxonomies, and providing realistic hands-on training without costly infrastructure. The resulting ITSs include: one employing exercises that scaffold comparison of conceptual categories, and another using lightweight simulations to provide experiential learning while circumventing typical cost and time overhead. Both systems incorporate instructional principles grounded in learning science research, including evidence-based features associated with ITS efficacy such as timely hints and feedback. Through iterative classroom deployment and refinement—including adding task-loop adaptivity to offer repeated practice until mastery—we observed encouraging learning outcomes, alongside insights into mitigating ''gaming the system'' behaviors. We detail our co-design process and formative evaluations—procedures, outcomes, and cautious interpretation due to the limited number of consented learners—and share lessons learned to inform scalable ITS development for cybersecurity workforce training in resource-constrained settings. Marshall An, Mahboobeh Mehrvarz, Leah Teffera, Matthew Kisow, Bruce M. McLaren, Christopher Bogart |
SIGCSE (1) | 1 |
| 2025 | Deceptive Overgeneralization in Adaptive Learning
Marshall An, John C. Stamper |
AIED (6) | 1 |
| 2024 | Singular Action, Complex Cognition: An Intelligent Tutoring System in Riichi Mahjong
Marshall An, Mufei He, John C. Stamper |
EC-TEL (2) | 1 |
| 2024 | Leveraging Intelligent Tutoring Systems to Enhance Project-Based Learning in Workforce Training at Community Colleges
Marshall An, Leah Teffera, Mahboobeh Mehrvarz, Bruce Li, Christopher Bogart, Majd F. Sakr, Bruce M. McLaren |
EC-TEL (2) | 1 |
| 2024 | What Factors Influence Persistence in Project-based Programming Courses at Community Colleges?abstractThe rapid adoption of emergent technologies is creating significant shortfall in the CS/IT workforce. With not enough students in the educational pipeline to meet the forthcoming demand over the next decade, community colleges are making the effort to train confident, knowledgeable, and self-driven workers in this field. Project-based learning (PBL) has been shown to be effective for these ends, but it poses distinct challenges in resource-limited community college contexts since it may require more time, preparation, and motivation than other teaching modalities, from both the student and the instructor. We studied fifteen sections of an introductory project-based Python course taught at six community colleges, investigating several features of PBL theorized to be particular barriers to student persistence, particularly among women and other identities traditionally underrepresented in technical fields. We describe successes and challenges faced by students in these areas and suggest implications for project-based learning curriculum and platform design. Christopher Bogart, Marshall An, Eric Keylor, Pawanjeet Singh, Jaromír Savelka, Majd F. Sakr |
SIGCSE (1) | 2 |
| 2024 | Programming Plagiarism Detection with Learner DataabstractCourses with programming assignments have long faced the issue of academic integrity violations (AIV) where cheating could harm the outcome of student learning. Checking code similarity in students' final submissions is a common way to mitigate this issue. But this single analysis is insufficient as 1) students can refactor their code to evade the check, 2) mere code similarity may not be strong enough evidence to support an AIV case, particularly for simpler assignments that may have similar solutions, and 3) code similarity cannot reveal much about the actual circumstances and behaviors of plagiarism. Due to the lack of supporting data or tools, many educators either abandon solving these challenges or rely on manual approaches that are not feasible at scale. In this paper, we propose a workflow to solve the above challenges for large programming classes by providing supporting evidence of cheating with additional learner data: detailed submission timelines with scores and source code. Running this workflow in a large advanced programming course over several years has helped us identify many cheating cases effectively and efficiently. Yifan Song 0007, Yuanxin Wang 0001, Marshall An, Christopher Bogart, Majd F. Sakr |
SIGCSE (2) | 3 |
| 2023 | Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming CoursesabstractThis paper studies recent developments in large language models’ (LLM) abilities to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. The emergence of ChatGPT resulted in heated debates of its potential uses (e.g., exercise generation, code explanation) as well as misuses in programming classes (e.g., cheating). Recent studies show that while the technology performs surprisingly well on diverse sets of assessment instruments employed in typical programming classes the performance is usually not sufficient to pass the courses. The release of GPT-4 largely emphasized notable improvements in the capabilities related to handling assessments originally designed for human test-takers. This study is the necessary analysis in the context of this ongoing transition towards mature generative AI systems. Specifically, we report the performance of GPT-4, comparing it to the previous generations of GPT models, on three Python courses with assessments ranging from simple multiple-choice questions (no code involved) to complex programming projects with code bases distributed into multiple files (599 exercises overall). Additionally, we analyze the assessments that were not handled well by GPT-4 to understand the current limitations of the model, as well as its capabilities to leverage feedback provided by an auto-grader. We found that the GPT models evolved from completely failing the typical programming class’ assessments (the original GPT-3) to confidently passing the courses with no human involvement (GPT-4). While we identified certain limitations in GPT-4’s handling of MCQs and coding exercises, the rate of improvement across the recent generations of GPT models strongly suggests their potential to handle almost any type of assessment widely used in higher education programming courses. These findings could be leveraged by educators and institutions to adapt the design of programming assessments as well as to fuel the necessary discussions into how programming classes should be updated to reflect the recent technological developments. This study provides evidence that programming instructors need to prepare for a world in which there is an easy-to-use widely accessible technology that can be utilized by learners to collect passing scores, with no effort whatsoever, on what today counts as viable programming knowledge and skills assessments. Jaromír Savelka, Arav Agarwal, Marshall An, Christopher Bogart, Majd F. Sakr |
ICER (1) | 3 |
| 2022 | Cheating Detection in Online Assessments via Timeline AnalysisabstractThe potential for academic integrity violations increases in online courses and instructors must place extra attention on academic integrity, since cheating techniques and costs are different than in the physical classroom. Although students are less supervised and able to study in a self-paced mode in online learning, unauthorized collaboration is still considered to be a serious integrity violation. However, online learning platforms have the advantage that they may capture detailed timelines of student activity. Analysis of these can enable instructors to detect many patterns of collaboration, e.g., working on assessments together, or copying solutions from unauthorized web pages. In this paper, we describe detection methods for several common patterns of alignment between work timelines of pairs of students, and these patterns' relationship with corroborative evidence such as similar answers and unusually fast completion times. We describe data collection necessary to apply the timeline analysis technique to weekly quiz assessments and project submissions, and discuss the strength of evidence the technique can provide in different situations. We have been applying these techniques in an online project-based course over several years, and it has helped instructors to successfully identify potential cheating cases. Jiameng Du, Yifan Song 0007, Mingxiao An, Marshall An, Christopher Bogart, Majd F. Sakr |
SIGCSE (1) | 4 |
| 2019 | An Intelligent-Agent Facilitated Scaffold for Fostering Reflection in a Team-Based Project Course
Sreecharan Sankaranarayanan, Xu Wang 0016, Cameron Dashti, Marshall An, Clarence Ngoh, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé |
AIED (2) | 4 |