Umair Z. Ahmed

dblp:18/11031 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Debugging and program repair · 48% Program verification · 18% Software testing · 17%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Computing education · 100%
Theoretical computer science
2 papers
Algorithms and data structures · 36% Automated reasoning and model checking · 27% Logic in computer science · 27%
Artificial intelligence
1 paper
Planning, search and constraint satisfaction · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
automated program repair
1.232022
Verifix: Verified Repair of Programming Assignments · ACM Trans. Softw. Eng. Methodol. 2022
Re-Factoring Based Program Repair Applied to Programming Assignments · ASE 2019
A feasibility study of using automated program repair for introductory programming assignments · ESEC/SIGSOFT FSE 2017
Computing education › programming education
automated feedback
0.832022
Targeted Example Generation for Compilation Errors · ASE 2019
A feasibility study of using automated program repair for introductory programming assignments · ESEC/SIGSOFT FSE 2017
Verifix: Verified Repair of Programming Assignments · ACM Trans. Softw. Eng. Methodol. 2022
Computing education
programming education
0.822020
Synthesizing Tasks for Block-based Programming · NeurIPS 2020
Targeted Example Generation for Compilation Errors · ASE 2019
Computing education › programming education
block-based programming
0.412020
Synthesizing Tasks for Block-based Programming · NeurIPS 2020
Software testing › mutation testing
program mutation
0.412020
Synthesizing Tasks for Block-based Programming · NeurIPS 2020
Program analysis
symbolic execution
0.412020
Synthesizing Tasks for Block-based Programming · NeurIPS 2020
Computing education
intelligent tutoring systems
0.312017
A feasibility study of using automated program repair for introductory programming assignments · ESEC/SIGSOFT FSE 2017
Debugging and program repair › automated program repair
repair of introductory programming assignments
0.312017
A feasibility study of using automated program repair for introductory programming assignments · ESEC/SIGSOFT FSE 2017
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
state space search
0.212015
Automatic Generation of Alternative Starting Positions for Simple Traditional Board Games · AAAI 2015
Logic in computer science › proof theory
natural deduction
0.212013
Automatically Generating Problems and Solutions for Natural Deduction · IJCAI 2013
Automated reasoning and model checking › theorem proving
proof generation
0.212013
Automatically Generating Problems and Solutions for Natural Deduction · IJCAI 2013
Computing education › programming education
programming assignment feedback
0.112019
Re-Factoring Based Program Repair Applied to Programming Assignments · ASE 2019
Program synthesis and code generation
example generation
0.112019
Targeted Example Generation for Compilation Errors · ASE 2019
Software testing
test generation
0.112019
Targeted Example Generation for Compilation Errors · ASE 2019
Computing education
educational technology
0.012013
Automatically Generating Problems and Solutions for Natural Deduction · IJCAI 2013
Computing education
problem generation
0.012013
Automatically Generating Problems and Solutions for Natural Deduction · IJCAI 2013

Methods — techniques the papers use, named apart from their topics

verification · 1.1predicate alignment · 1.1symbolic execution · 0.9monte carlo tree search · 0.9supervised classification · 0.8specification inference · 0.8search-based synthesis · 0.8refactoring · 0.8dense neural network · 0.8MaxSMT · 0.6Max-SMT · 0.6symbolic methods · 0.4iterative simulation · 0.4
YearPublicationVenuePosition
2025 Feasibility Study of Augmenting Teaching Assistants with AI for CS1 Programming Feedback
abstract
With the increasing adoption of Large Language Models (LLMs), there are proposals to replace human Teaching Assistants (TAs) with LLM-based AI agents for providing feedback to students. In this paper, we explore a new hybrid model where human TAs receive AI-generated feedback for CS1 programming exercises, which they can then review and modify as needed. We conducted a large-scale randomized intervention with 185 CS1 undergraduate students, comparing the efficacy of this hybrid approach against manual feedback and direct AI-generated feedback.
Umair Z. Ahmed, Shubham Sahai, Ben Leong, Amey Karkare
SIGCSE (1)1
2022 Verifix: Verified Repair of Programming Assignments
abstract
Automated feedback generation for introductory programming assignments is useful for programming education. Most works try to generate feedback to correct a student program by comparing its behavior with an instructor’s reference program on selected tests. In this work, our aim is to generate verifiably correct program repairs as student feedback. A student-submitted program is aligned and composed with a reference solution in terms of control flow, and the variables of the two programs are automatically aligned via predicates describing the relationship between the variables. When verification attempt for the obtained aligned program fails, we turn a verification problem into a MaxSMT problem whose solution leads to a minimal repair. We have conducted experiments on student assignments curated from a widely deployed intelligent tutoring system. Our results show that generating verified repair without sacrificing the overall repair rate is possible. In fact, our implementation, Verifix, is shown to outperform Clara, a state-of-the-art tool, in terms of repair rate. This shows the promise of using verified repair to generate high confidence feedback in programming pedagogy settings.
Umair Z. Ahmed, Zhiyu Fan, Jooyong Yi, Omar I. Al-Bataineh, Abhik Roychoudhury
ACM Trans. Softw. Eng. Methodol.1
2020 MACER: A Modular Framework for Accelerated Compilation Error Repair
Darshak Chhatbar, Umair Z. Ahmed, Purushottam Kar
AIED (1)2
2020 Synthesizing Tasks for Block-based Programming
abstract
Block-based visual programming environments play a critical role in introducing computing concepts to K-12 students. One of the key pedagogical challenges in these environments is in designing new practice tasks for a student that match a desired level of difficulty and exercise specific programming concepts. In this paper, we formalize the problem of synthesizing visual programming tasks. In particular, given a reference visual task $\task^{in}$ and its solution code $\code^{in}$, we propose a novel methodology to automatically generate a set $\{(\task^{out}, \code^{out})\}$ of new tasks along with solution codes such that tasks $\task^{in}$ and $\task^{out}$ are conceptually similar but visually dissimilar. Our methodology is based on the realization that the mapping from the space of visual tasks to their solution codes is highly discontinuous; hence, directly mutating reference task $\task^{in}$ to generate new tasks is futile. Our task synthesis algorithm operates by first mutating code $\code^{in}$ to obtain a set of codes $\{\code^{out}\}$. Then, the algorithm performs symbolic execution over a code $\code^{out}$ to obtain a visual task $\task^{out}$; this step uses the Monte Carlo Tree Search (MCTS) procedure to guide the search in the symbolic tree. We demonstrate the effectiveness of our algorithm through an extensive empirical evaluation and user study on reference tasks taken from the Hour of Code: Classic Maze challenge by Code.org and the Intro to Programming with Karel course by CodeHS.com.
Umair Z. Ahmed, Maria Christakis, Aleksandr Efremov, Nigel Fernandez, Ahana Ghosh, Abhik Roychoudhury, Adish Singla
NeurIPS1
2019 Targeted Example Generation for Compilation Errors
abstract
We present TEGCER, an automated feedback tool for novice programmers. TEGCER uses supervised classification to match compilation errors in new code submissions with relevant pre-existing errors, submitted by other students before. The dense neural network used to perform this classification task is trained on 15000+ error-repair code examples. The proposed model yields a test set classification Pred@3 accuracy of 97.7% across 212 error category labels. Using this model as its base, TEGCER presents students with the closest relevant examples of solutions for their specific error on demand. A large scale (N>230) usability study shows that students who use TEGCER are able to resolve errors more than 25% faster on average than students being assisted by human tutors.
Umair Z. Ahmed, Renuka Sindhgatta, Nisheeth Srivastava, Amey Karkare
ASE1
2019 Re-Factoring Based Program Repair Applied to Programming Assignments
abstract
Automated program repair has been used to provide feedback for incorrect student programming assignments, since program repair captures the code modification needed to make a given buggy program pass a given test-suite. Existing student feedback generation techniques are limited because they either require manual effort in the form of providing an error model, or require a large number of correct student submissions to learn from, or suffer from lack of scalability and accuracy. In this work, we propose a fully automated approach for generating student program repairs in real-time. This is achieved by first re-factoring all available correct solutions to semantically equivalent solutions. Given an incorrect program, we match the program with the closest matching refactored program based on its control flow structure. Subsequently, we infer the input-output specifications of the incorrect program's basic blocks from the executions of the correct program's aligned basic blocks. Finally, these specifications are used to modify the blocks of the incorrect program via search-based synthesis. Our dataset consists of almost 1,800 real-life incorrect Python program submissions from 361 students for an introductory programming course at a large public university. Our experimental results suggest that our method is more effective and efficient than recently proposed feedback generation approaches. About 30% of the patches produced by our tool Refactory are smaller than those produced by the state-of-art tool Clara, and can be produced given fewer correct solutions (often a single correct solution) and in a shorter time. We opine that our method is applicable not only to programming assignments, and could be seen as a general-purpose program repair method that can achieve good results with just a single correct reference solution.
Umair Z. Ahmed, Sergey Mechtaev, Ben Leong, Abhik Roychoudhury
ASE2
2017 A feasibility study of using automated program repair for introductory programming assignments
abstract
Despite the fact an intelligent tutoring system for programming (ITSP) education has long attracted interest, its widespread use has been hindered by the difficulty of generating personalized feedback automatically. Meanwhile, automated program repair (APR) is an emerging new technology that automatically fixes software bugs, and it has been shown that APR can fix the bugs of large real-world software. In this paper, we study the feasibility of marrying intelligent programming tutoring and APR. We perform our feasibility study with four state-of-the-art APR tools (GenProg, AE, Angelix, and Prophet), and 661 programs written by the students taking an introductory programming course. We found that when APR tools are used out of the box, only about 30% of the programs in our dataset are repaired. This low repair rate is largely due to the student programs often being significantly incorrect - in contrast, professional software for which APR was successfully applied typically fails only a small portion of tests. To bridge this gap, we adopt in APR a new repair policy akin to the hint generation policy employed in the existing ITSP. This new repair policy admits partial repairs that address part of failing tests, which results in 84% improvement of repair rate. We also performed a user study with 263 novice students and 37 graders, and identified an understudied problem; while novice students do not seem to know how to effectively make use of generated repairs as hints, the graders do seem to gain benefits from repairs.
Jooyong Yi, Umair Z. Ahmed, Amey Karkare, Shin Hwei Tan, Abhik Roychoudhury
ESEC/SIGSOFT FSE2
2015 Automatic Generation of Alternative Starting Positions for Simple Traditional Board Games
abstract
Simple board games, like Tic-Tac-Toe and CONNECT-4, play an important role not only in the development of mathematical and logical skills, but also in the emotional and social development. In this paper, we address the problem of generating targeted starting positions for such games. This can facilitate new approaches for bringing novice players to mastery, and also leads to discovery of interesting game variants. We present an approach that generates starting states of varying hardness levels for player 1 in a two-player board game, given rules of the board game, the desired number of steps required for player 1 to win, and the expertise levels of the two players. Our approach leverages symbolic methods and iterative simulation to efficiently search the extremely large state space. We present experimental results that include discovery of states of varying hardness levels for several simple grid-based board games. The presence of such states for standard game variants like 4 x 4 Tic-Tac-Toe opens up new games to be played that have never been played as the default start state is heavily biased.
Umair Z. Ahmed, Krishnendu Chatterjee, Sumit Gulwani
AAAI1
2013 Automatically Generating Problems and Solutions for Natural Deduction
Umair Z. Ahmed, Sumit Gulwani, Amey Karkare
IJCAI1
2012 Can Modern Statistical Parsers Lead to Better Natural Language Understanding for Education?
Umair Z. Ahmed, Arpit Kumar, Monojit Choudhury, Kalika Bali
CICLing (1)1