VLDB 2026 Research / reviewers in the wild / expert
Muntasir Hoq
dblp:355/9928
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0003-2591-0476ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Explainable AI Assistant for Introductory Programming Education: Improving Feedback Reliability with Instructor-AI Collaboration
Muntasir Hoq, Griffin Pitts, Bradford W. Mott, Seung Y. Lee, Jessica Vandenberg, Shuyin Jiao, Narges Norouzi, James C. Lester, Bita Akram |
AIED (1) | 1 |
| 2026 | Personalized Worked Example Generation from Student Code Submissions Using Pattern-based Knowledge ComponentsabstractAdaptive programming practice often relies on fixed libraries of worked examples and practice problems, which require substantial authoring effort and may not correspond well to the logical errors and partial solutions students produce while writing code. As a result, students may receive learning content that does not directly address the concepts they are working to understand, while instructors must either invest additional effort in expanding content libraries or accept a coarse level of personalization. We present an approach for knowledge-component (KC) guided educational content generation using pattern-based KCs extracted from student code. Given a problem statement and student submissions, our pipeline extracts recurring structural KC patterns from students' code through AST-based analysis and uses them to condition a generative model. In this study, we apply this approach to worked example generation, and compare baseline and KC-conditioned outputs through expert evaluation. Results suggest that KC-conditioned generation improves topical focus and relevance to students' underlying logical errors, providing evidence that KC-based steering of generative models can support personalized learning at scale. Griffin Pitts, Muntasir Hoq, Peter Brusilovsky, Narges Norouzi, Arto Hellas, Juho Leinonen 0001, Bita Akram |
L@S | 2 |
| 2026 | INSIGHT: An Explainable, Instructor-Guided AI Assistant for Active Learning in CS1abstractActive learning in introductory programming depends on frequent, high-quality feedback, yet instructors often struggle to deliver consistent support at scale. In this poster, we introduce INSIGHT, an AI-driven classroom assistant designed to promote active learning in introductory programming (CS1) courses through scalable, personalized, and explainable feedback. The assistant combines the generative capabilities of large language models (LLMs) with instructor-in-the-loop authoring and an explainable code analysis engine to ensure pedagogically aligned support. Instructors can co-design problems with LLM assistance, provide exemplar solutions, define common student errors, and author targeted feedback. The AI engine analyzes student code submissions, identifies misconceptions, and maps them to instructor-verified feedback in real time. INSIGHT is designed to ensure that key educational concepts and common misconceptions are explicitly addressed by instructors, while also leveraging LLMs to provide reasonable feedback for novel or edge-case solutions that instructors may not have anticipated. By combining instructor expertise with the flexibility of generative AI, the assistant helps close feedback gaps and ensures more comprehensive coverage of student learning needs, especially in large or diverse classrooms with limited instructional support. Muntasir Hoq, Jessica Vandenberg, Seung Y. Lee, Bradford W. Mott, James C. Lester, Narges Norouzi, Shuyin Jiao, Bita Akram |
SIGCSE (2) | 1 |
| 2026 | Automated Program Repair of Uncompilable Student CodeabstractA significant portion of student programming submissions in CS1 learning environments are uncompilable, limiting their use in student modeling and downstream knowledge tracing. Traditional modeling pipelines often exclude these cases, discarding observations of student learning. This study investigates automated program repair as a strategy to recover uncompilable code while preserving students' structural intent for use in student modeling. Within this framework, we assess large language models (LLMs) as repair agents under high- and low-context prompting conditions. Repairs were evaluated for compilability, edit distance, and preservation of students' original structure and logic. While all models produced compilable repairs, they differed in how well they preserve students' control flow and code structure, affecting their pedagogical utility. By recovering uncompilable submissions, this work enables richer and more comprehensive analyses of learners' coding processes and development over time. Griffin Pitts, Aum Pandya, Darsh Rank, Muntasir Hoq, Tirth Bhatt, Bita Akram |
SIGCSE (2) | 4 |
| 2025 | Automated Identification of Logical Errors in Programs: Advancing Scalable Analysis of Student Misconceptions
Muntasir Hoq, Ananya Rao, Reisha Jaishankar, Krish Piryani, Nithya Janapati, Jessica Vandenberg, Bradford W. Mott, Narges Norouzi, James C. Lester, Bita Akram |
EDM | 1 |
| 2025 | Privacy-Preserving Distributed Link Predictions Among Peers in Online Classrooms Using Federated Learning
Anurata Prabha Hridi, Muntasir Hoq, Zhikai Gao, Collin F. Lynch, Rajeev Sahay, Seyyedali Hosseinalipour, Bita Akram |
EDM | 2 |
| 2025 | An Automated Approach to Recommending Relevant Worked Examples for Programming ProblemsabstractNovice programmers can greatly improve their understanding of challenging programming concepts by studying worked examples that demonstrate the implementation of these concepts. Despite the extensive repositories of effective worked examples created by CS education experts, a key challenge remains: identifying the most relevant worked example for a given programming problem and the specific difficulties a student faces solving the problem. Previous studies have explored similar example recommendation approaches. Our research introduces a novel method by utilizing deep learning code representation models to generate code vectors, capturing both syntactic and semantic similarities among programming examples. Driven by the need to provide relevant and personalized examples to programming students, our approach emphasizes similarity assessment and clustering techniques to identify similar code problems, examples, and challenges. This method aims to deliver more accurate and contextually relevant recommendations based on individual learning needs. Providing tailored support to students in real-time facilitates better problem-solving strategies and enhances students' learning experiences, contributing to the advancement of programming education. Muntasir Hoq, Atharva Patil, Kamil Akhuseyinoglu, Peter Brusilovsky, Bita Akram |
SIGCSE (1) | 1 |
| 2024 | Detecting ChatGPT-Generated Code Submissions in a CS1 Course Using Machine Learning ModelsabstractThe emergence of publicly accessible large language models (LLMs) such as ChatGPT poses unprecedented risks of new types of plagiarism and cheating where students use LLMs to solve exercises for them. Detecting this behavior will be a necessary component in introductory computer science (CS1) courses, and educators should be well-equipped with detection tools when the need arises. However, ChatGPT generates code non-deterministically, and thus, traditional similarity detectors might not suffice to detect AI-created code. In this work, we explore the affordances of Machine Learning (ML) models for the detection task. We used an openly available dataset of student programs for CS1 assignments and had ChatGPT generate code for the same assignments, and then evaluated the performance of both traditional machine learning models and Abstract Syntax Tree-based (AST-based) deep learning models in detecting ChatGPT code from student code submissions. Our results suggest that both traditional machine learning models and AST-based deep learning models are effective in identifying ChatGPT-generated code with accuracy above 90%. Since the deployment of such models requires ML knowledge and resources that are not always accessible to instructors, we also explore the patterns detected by deep learning models that indicate possible ChatGPT code signatures, which instructors could possibly use to detect LLM-based cheating manually. We also explore whether explicitly asking ChatGPT to impersonate a novice programmer affects the code produced. We further discuss the potential applications of our proposed models for enhancing introductory computer science instruction. Muntasir Hoq, Yang Shi 0004, Juho Leinonen 0001, Damilola Babalola, Collin F. Lynch, Thomas W. Price, Bita Akram |
SIGCSE (1) | 1 |
| 2024 | Towards Attention-Based Automatic Misconception Identification in Introductory Programming CoursesabstractIdentifying misconceptions in student programming solutions is an important step in evaluating their comprehension of fundamental programming concepts. While misconceptions are latent constructs that are hard to evaluate directly from student programs, logical errors can signal their existence in students' understanding. Tracing multiple occurrences of related logical bugs over different problems can provide strong evidence of students' misconceptions. This study presents preliminary results of utilizing an interpretable state-of-the-art Abstract Syntax Tree-based embedding neural network to identify logical mistakes in student code. In this study, we show a proof-of-concept of the errors identified in student programs by classifying correct versus incorrect programs. Our preliminary results show that our framework is able to automatically identify misconceptions without designing and applying a detailed rubric. This approach shows promise for improving the quality of instruction in introductory programming courses by providing educators with a powerful tool that offers personalized feedback while enabling accurate modeling of student misconceptions. Muntasir Hoq, Jessica Vandenberg, Bradford W. Mott, James C. Lester, Narges Norouzi, Bita Akram |
SIGCSE (2) | 1 |
| 2024 | Use of Large Language Models for Extracting Knowledge Components in CS1 Programming ExercisesabstractThis study utilizes large language models to extract foundational programming concepts in programming assignments in a CS1 course. We seek to answer the following research questions: RQ1. How effectively can large language models identify knowledge components in a CS1 course from programming assignments? RQ2. Can large language models be used to extract program-level knowledge components, and how can the information be used to identify students' misconceptions? Preliminary results demonstrated a high similarity between course-level knowledge components retrieved from a large language model and that of an expert-generated list. Rose Niousha, Muntasir Hoq, Bita Akram, Narges Norouzi |
SIGCSE (2) | 2 |
| 2023 | SANN: Programming Code Representation Using Attention Neural Network with Optimized Subtree ExtractionabstractAutomated analysis of programming data using code representation methods offers valuable services for programmers, from code completion to clone detection to bug detection. Recent studies show the effectiveness of Abstract Syntax Trees (AST), pre-trained Transformer-based models, and graph-based embeddings in programming code representation. However, pre-trained large language models lack interpretability, while other embedding-based approaches struggle with extracting important information from large ASTs. This study proposes a novel Subtree-based Attention Neural Network (SANN) to address these gaps by integrating different components: an optimized sequential subtree extraction process using Genetic algorithm optimization, a two-way embedding approach, and an attention network. We investigate the effectiveness of SANN by applying it to two different tasks: program correctness prediction and algorithm detection on two educational datasets containing both small and large-scale code snippets written in Java and C, respectively. The experimental results show SANN's competitive performance against baseline models from the literature, including code2vec, ASTNN, TBCNN, CodeBERT, GPT-2, and MVG, regarding accurate predictive power. Finally, a case study is presented to show the interpretability of our model prediction and its application for an important human-centered computing application, student modeling. Our results indicate the effectiveness of the SANN model in capturing important syntactic and semantic information from students' code, allowing the construction of accurate student models, which serve as the foundation for generating adaptive instructional support such as individualized hints and feedback. Muntasir Hoq, Sushanth Reddy Chilla, Melika Ahmadi Ranjbar, Peter Brusilovsky, Bita Akram |
CIKM | 1 |
| 2023 | Analysis of an Explainable Student Performance Prediction Model in an Introductory Programming Course
Muntasir Hoq, Peter Brusilovsky, Bita Akram |
EDM | 1 |