Abbas Heydarnoori

dblp:h/AbbasHeydarnoori · DBLP profile ↗
← Back
32ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0001-9785-2880ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 29 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 An Empirical Analysis of Test Failures in AI-Generated Pull Requests
abstract
Recent advances in large language models (LLMs) and AI agents have enabled automated test generation to become a practical component of modern software development workflows. However, these systems frequently produce tests that fail to compile or execute correctly, raising questions about their robustness in real-world settings. In this study, we conduct an empirical analysis of test failures in the AIDev dataset from the MSR 2026 Mining Challenge. Particularly, we examine 500 merged pull requests containing test-related failures (100 per agent across 5 AI coding agents) and classify each case using a refined taxonomy that distinguishes between compile-time and runtime errors. Our findings reveal that runtime errors dominate (62.6% vs 37.4% compile-time), with assertion failures being the most common issue (28.6%). The analysis exposes distinct failure profiles across agents and provides insights for improving AI-driven testing reliability in large-scale development environments. Our replication package is available at: https://figshare.com/s/005a3942818fba893a7e.
Alireza Hoseinpour, Sajjad Rezvani Boroujeni, Jashhvanth Tamilselvan Kunthavai, Kyle Cusimano, Abbas Heydarnoori
MSR5
2026 Readability of AI-Generated Pull Request Descriptions Across Pull Request Types
abstract
This study investigates the readability of AI-generated pull request (PR) descriptions, a critical factor for enabling programmers to efficiently understand and integrate code changes. We analyze a large-scale and diverse dataset, AIDev, which contains AI-generated PR descriptions spanning multiple PR types, including bug fix, feature, documentation, test, dependency, and refactor changes. Readability is measured using the Flesch Reading Ease metric, complemented by a quantitative rubric-based evaluation to assess completeness. We compare readability and completeness across PR types and AI agents. Our results reveal systematic differences in both readability and completeness that are influenced by the PR type as well as the underlying AI model. These findings provide empirical evidence of current AI models’ performance in generating PR documentation, highlight limitations in existing approaches, and identify opportunities to improve the design of AI-assisted tools to support clearer, more complete, and more effective developer workflows.
Aidan Tobar, Joseph Peterson, Abbas Heydarnoori
MSR3
2025 Enhancing Commit Message Generation in Software Repositories: A RAG-Based Approach
abstract
Creating clear and detailed commit messages manually is both time-consuming and prone to inconsistency. Existing automated methods, such as rule-based templates, retrieval-based systems, and neural sequence-to-sequence models, often fail to capture the full context of code changes, resulting in incomplete or inaccurate messages. To overcome these limitations, we introduce a Retrieval-Augmented Generation (RAG) framework that combines a pretrained retriever with transformer-based generators (BART and T5). Our approach builds a comprehensive knowledge base of code diffs and commit messages, incorporating relevant external context through joint end-to-end training. The dataset was sourced from the top GitHub repositories and curated using three key filtering and cleaning steps: (1) excluding non-target programming languages via regular expressions, (2) detecting the presence of “What” and “Why” information in commit messages with a Bidirectional LSTM model, and (3) manually evaluating high-quality messages with expert assistance. Experiments on curated datasets in Python, JavaScript, and Java demonstrate that our BART-based configuration significantly improves BLEU, METEOR, and ROUGE scores compared to existing methods, with the RAG-based models consistently outperforming other approaches across all evaluation metrics.
Mehrdad Yadollahi, Abbas Heydarnoori, Iman Khazrak
SERA2
2025 Predicting the understandability of computational notebooks through code metrics analysis
Mojtaba Mostafavi Ghahfarokhi, Alireza Asadi, Arash Asgari, Bardia Mohammadi, Abbas Heydarnoori
Empir. Softw. Eng.5
2024 A Roadmap for Enriching Jupyter Notebooks Documentation with Kaggle Data
abstract
Recent advancements in AI and data science have led to the increased use of Jupyter notebooks. As such, various AI-Based automated tools have been also developed to automatically document notebooks. However, a key challenge is the absence of suitable datasets for training AI models. In this paper, we outline a valuable roadmap for developing a dataset of (markdown, code) pairs centered on functions in Jupyter notebooks. The roadmap encompasses four high-level steps: structural filtering, structural processing, conceptual filtering, and conceptual processing. Our proposed roadmap leads to providing a quality dataset for training AI models on Jupyter notebooks.
Mojtaba Mostafavi Ghahfarokhi, Hamed Jahantigh, Alireza Asadi, Sepehr Kianiangolafshani, Ashkan Khademian, Abbas Heydarnoori
CAIN6
2024 Beyond Syntax: Unleashing the Power of Computational Notebooks Code Metrics in Documentation Generation
abstract
Computational notebooks, like Kaggle notebooks, offer an integrated platform for coding and documentation, yet the latter's quality often falls short as scientists may neglect this crucial aspect. This paper addresses the need for improved and efficient code documentation generation in computational notebooks. As recent literature emphasizes integrating code's inherent structure into documentation generation models, our research explores unutilized structural characteristics, incorporating metrics from code sequences to enable better code documentation suggestions. Evidenced by the improved BLEU scores, our proposed method significantly outperforms the conventional model in a preliminary 10-fold cross-validation experiment and further provides a flexible foundation for integrating source code metrics into diverse code documentation generation models.
Mojtaba Mostafavi Ghahfarokhi, Ashkan Khademian, Sepehr Kianiangolafshani, Alireza Asadi, Hamed Jahantigh, Abbas Heydarnoori
CAIN6
2024 Can Code Metrics Enhance Documentation Generation for Computational Notebooks?
abstract
In software development, code documentation is crucial for collaboration and maintenance, especially as projects become more complex. However, it is often neglected due to the tedious effort it requires. This paper explores automating documentation generation for computational notebooks, focusing on the impact of code metrics such as lines of code, API popularity, and complexity on this task. Using a dataset of 22K code-documentation pairs, we compare deep learning models with and without code metric augmentation. The results show that incorporating these metrics significantly improves the accuracy of documentation generation, underscoring the connection between code metrics and quality documentation.
Mojtaba Mostafavi Ghahfarokhi, Hamed Jahantigh, Sepehr Kianiangolafshani, Ashkan Khademian, Alireza Asadi, Abbas Heydarnoori
ASE6
2024 DATAR: A Dataset for Tracking App Releases
abstract
Android apps continuously evolve to meet user expectations and thrive in the competitive environment of app stores. Hence, making informed decisions is crucial for the success of upcoming releases. In recent years, researchers have sought to aid developers in release planning by studying, analyzing, and modeling evolutionary information derived from tracking releases of various apps. They have demonstrated how the types of information provided on different platforms can effectively predict and assess the quality or success of an app. Nevertheless, this field needs a comprehensive dataset containing evolutionary information that tracks various releases of open-source Android apps. Existing datasets lack release-level information, are not publicly available, or do not cover a wide range of up-to-date apps. This paper introduces a dataset comprising diverse metadata adapted from GitHub and Google Play, along with impactful metrics extracted from the source code of 8,041 published releases of 1,363 open-source Android apps.
Yasaman Abedini, Mohammad Hadi Hajihosseini, Abbas Heydarnoori
MSR3
2024 DistilKaggle: A Distilled Dataset of Kaggle Jupyter Notebooks
abstract
Jupyter notebooks have become indispensable tools for data analysis and processing in various domains. However, despite their widespread use, there is a notable research gap in understanding and analyzing the contents and code metrics of these notebooks. This gap is primarily attributed to the absence of datasets that encompass both Jupyter notebooks and extracted their code metrics. To address this limitation, we introduce DistilKaggle, a unique dataset specifically curated to facilitate research on code metrics in Jupyter notebooks, utilizing the Kaggle repository as a prime source. Through an extensive study, we identify thirty-four code metrics that significantly impact Jupyter notebook code quality. These features such as lines of code cell, mean number of words in markdown cells, performance tier of developer, etc., are crucial for understanding and improving the overall effectiveness of computational notebooks. The DistilKaggle dataset which is derived from a vast collection of notebooks constitutes two distinct datasets: (i) Code Cells and Markdown Cells Dataset which is presented in two CSV files, allowing for easy integration into researchers' workflows as dataframes. It provides a granular view of the content structure within 542,051 Jupyter notebooks, enabling detailed analysis of code and markdown cells; and (ii) The Notebook Code Metrics Dataset focused on the identified code metrics of notebooks. Researchers can leverage this dataset to access Jupyter notebooks with specific code quality characteristics, surpassing the limitations of filters available on the Kaggle website. Furthermore, the reproducibility of the notebooks in our dataset is ensured through the code cells and markdown cells datasets, offering a reliable foundation for researchers to build upon. Given the substantial size of our datasets, it becomes an invaluable resource for the research community, surpassing the capabilities of individual Kaggle users to collect such extensive data. For accessibility and transparency, both the dataset and the code utilized in crafting this dataset are publicly available at https://github.com/ISE-Research/DistilKaggle.
Mojtaba Mostafavi Ghahfarokhi, Arash Asgari, Mohammad Abolnejadian, Abbas Heydarnoori
MSR4
2024 GIRT-Model: Automated Generation of Issue Report Templates
abstract
Platforms such as GitHub and GitLab introduce Issue Report Templates (IRTs) to enable more effective issue management and better alignment with developer expectations. However, these templates are not widely adopted in most repositories, and there is currently no tool available to aid developers in generating them. In this work, we introduce GIRT-Model, an assistant language model that automatically generates IRTs based on the developer's instructions regarding the structure and necessary fields. We create GIRT-Instruct, a dataset comprising pairs of instructions and IRTs, with the IRTs sourced from GitHub repositories. We use GIRT-Instruct to instruction-tune a T5-base model to create the GIRT-Model.
Nafiseh Nikeghbal, Amir Hossein Kargaran, Abbas Heydarnoori
MSR3
2024 Can GitHub Issues Help in App Review Classifications?
abstract
App reviews reflect various user requirements that can aid in planning maintenance tasks. Recently, proposed approaches for automatically classifying user reviews rely on machine learning algorithms. A previous study demonstrated that models trained on existing labeled datasets exhibit poor performance when predicting new ones. Therefore, a comprehensive labeled dataset is essential to train a more precise model. In this paper, we propose a novel approach that assists in augmenting labeled datasets by utilizing information extracted from an additional source, GitHub issues, that contains valuable information about user requirements. First, we identify issues concerning review intentions (bug reports, feature requests, and others) by examining the issue labels. Then, we analyze issue bodies and define 19 language patterns for extracting targeted information. Finally, we augment the manually labeled review dataset with a subset of processed issues through the Within-App , Within-Context , and Between-App Analysis methods. We conducted several experiments to evaluate the proposed approach. Our results demonstrate that using labeled issues for data augmentation can improve the F1-score to 6.3 in bug reports and 7.2 in feature requests. Furthermore, we identify an effective range of 0.3 to 0.7 for the auxiliary volume, which provides better performance improvements.
Yasaman Abedini, Abbas Heydarnoori
ACM Trans. Softw. Eng. Methodol.2
2023 GIRT-Data: Sampling GitHub Issue Report Templates
abstract
GitHub’s issue reports provide developers with valuable information that is essential to the evolution of a software development project. Contributors can use these reports to perform software engineering tasks like submitting bugs, requesting features, and collaborating on ideas. In the initial versions of issue reports, there was no standard way of using them. As a result, the quality of issue reports varied widely. To improve the quality of issue reports, GitHub introduced issue report templates (IRTs), which pre-fill issue descriptions when a new issue is opened. An IRT usually contains greeting contributors, describing project guidelines, and collecting relevant information. However, despite of effectiveness of this feature which was introduced in 2016, only nearly 5% of GitHub repositories (with more than 10 stars) utilize it. There are currently few articles on IRTs, and the available ones only consider a small number of repositories.In this work, we introduce GIRT-DATA, the first and largest dataset of IRTs in both YAML and Markdown format. This dataset and its corresponding open-source crawler tool are intended to support research in this area and to encourage more developers to use IRTs in their repositories. The stable version of the dataset contains 1,084,300 repositories and 50,032 of them support IRTs. The stable version of the dataset and crawler is available here: https://github.com/kargaranamir/girt-data
Nafiseh Nikeghbal, Amir Hossein Kargaran, Abbas Heydarnoori, Hinrich Schütze
MSR3
2023 Semantically-enhanced topic recommendation systems for software projects
Maliheh Izadi, Mahtab Nejati, Abbas Heydarnoori
Empir. Softw. Eng.3
2022 The ineffectiveness of domain-specific word embedding models for GUI test reuse
abstract
Reusing test cases across similar applications can significantly reduce testing effort. Some recent test reuse approaches successfully exploit word embedding models to semantically match GUI events across Android apps. It is a common understanding that word embedding models trained on domain-specific corpora perform better on specialized tasks. Our recent study confirms this understanding in the context of Android test reuse. It shows that word embedding models trained with a corpus of the English descriptions of apps in the Google Play Store lead to a better semantic matching of Android GUI events. Motivated by this result, we hypothesize that we can further increase the effectiveness of semantic matching by partitioning the corpus of app descriptions into domain-specific corpora. Our experiments do not confirm our hypothesis. This paper sheds light on this unexpected negative result that contradicts the common understanding.
Farideh Khalili, Ali Mohebbi 0003, Valerio Terragni, Mauro Pezzè, Leonardo Mariani, Abbas Heydarnoori
ICPC6
2022 Predicting the objective and priority of issue reports in software repositories
Maliheh Izadi, Kiana Akbari, Abbas Heydarnoori
Empir. Softw. Eng.3
2021 Automated Recovery of Issue-Commit Links Leveraging Both Textual and Non-textual Data
abstract
An issue report documents the discussions around required changes in issue-tracking systems, while a commit contains the change itself in the version control systems. Recovering links between issues and commits can facilitate many software evolution tasks such as bug localization, defect prediction, software quality measurement, and software documentation. A previous study on over half a million issues from GitHub reports only about 42.2% of issues are manually linked by developers to their pertinent commits. Automating the linking of commit-issue pairs can contribute to the improvement of the said tasks. By far, current state-of-the-art approaches for automated commit-issue linking suffer from low precision, leading to unreliable results, sometimes to the point that imposes human supervision on the predicted links. The low performance gets even more severe when there is a lack of textual information in either commits or issues. Current approaches are also proven computationally expensive. We propose Hybrid-Linker, an enhanced approach that overcomes such limitations by exploiting two information channels; (1) a non-textual-based component that operates on non-textual, automatically recorded information of the commit-issue pairs to predict a link, and (2) a textual-based one which does the same using textual information of the commit-issue pairs. Then, combining the results from the two classifiers, Hybrid-Linker makes the final prediction. Thus, every time one component falls short in predicting a link, the other component fills the gap and improves the results. We evaluate Hybrid-Linker against competing approaches, namely FRLink and DeepLink on a dataset of 12 projects. Hybrid-Linker achieves 90.1%, 87.8%, and 88.9% based on recall, precision, and F-measure, respectively. It also outperforms FRLink and DeepLink by 31.3%, and 41.3%, regarding the F-measure. Moreover, the proposed approach exhibits extensive improvements in terms of performance as well. Finally, our source code and data are publicly available.
Pooya Rostami Mazrae, Maliheh Izadi, Abbas Heydarnoori
ICSME3
2021 Topic recommendation for software repositories using multi-label classification algorithms
Maliheh Izadi, Abbas Heydarnoori, Georgios Gousios
Empir. Softw. Eng.2
2020 Improving Quality of a Post's Set of Answers in Stack Overflow
abstract
Community Question Answering platforms such as Stack Overflow help a wide range of users solve their challenges on-line. As the popularity of these communities has grown over the years, both the number of members and posts have escalated. Also, due to the diverse backgrounds, skills, expertise, and viewpoints of users, each question may obtain more than one answer. Therefore, the focus has changed toward producing posts that have a set of answers more valuable for the community as a whole, not just one accepted-answer aimed at satisfying only the question-asker. Same as every universal community, a large number of low-quality posts on Stack Overflow require improvement. We call these posts "deficient", and define them as posts with questions that either have no answer yet or can be improved by other ones. In this paper, we propose an approach to automate the identification process of such posts and boost their set of answers, utilizing the help of related experts. With the help of 60 participants, we trained a classification model to identify deficient posts by investigating the relationship between characteristics of 3075 questions posted on Stack Overflow and their need for better answers set. Then, we developed an Eclipse plugin named SOPI and integrated the prediction model in the plugin to link these deficient posts to related developers (in terms of their development context and expertise area) and help them improve the answer set. We evaluated both the functionality of our plugin and the impact of answers submitted to Stack Overflow with the help of 10 and 15 expert industrial developers, respectively. Our results indicate that decision trees, specifically the J48 algorithm, predicts a deficient question better than the other methods with 94.5% precision and 90.3% recall. We conclude that not only our plugin helps programmers contribute more easily to Stack Overflow, but also it improves the quality of existing answers.
MohammadReza Tavakoli, Maliheh Izadi, Abbas Heydarnoori
SEAA3
2020 Generating summaries for methods of event-driven programs: An Android case study
Alireza Aghamohammadi, Maliheh Izadi, Abbas Heydarnoori
J. Syst. Softw.3
2020 Studying the Relationship Between the Usage of APIs Discussed in the Crowd and Post-Release Defects
Hamed Tahmooresi, Abbas Heydarnoori, Reza Nadri
J. Syst. Softw.2
2019 Cross-project code clones in GitHub
Mohammad Gharehyazie, Baishakhi Ray, Mehdi Keshani, Masoumeh Soleimani Zavosht, Abbas Heydarnoori, Vladimir Filkov
Empir. Softw. Eng.5
2018 Microservices migration patterns
abstract
Summary Microservices architectures are becoming the defacto standard for building continuously deployed systems. At the same time, there is a substantial growth in the demand for migrating on‐premise legacy applications to the cloud. In this context, organizations tend to migrate their traditional architectures into cloud‐native architectures using microservices. This article reports a set of migration and rearchitecting design patterns that we have empirically identified and collected from industrial‐scale software migration projects. These migration patterns can help information technology organizations plan their migration projects toward microservices more efficiently and effectively. In addition, the proposed patterns facilitate the definition of migration plans by pattern composition. Qualitative empirical research is used to evaluate the validity of the proposed patterns. Our findings suggest that the proposed patterns are evident in other architectural refactoring and migration projects and strong candidates for effective patterns in system migrations.
Armin Balalaie, Abbas Heydarnoori, Pooyan Jamshidi, Damian A. Tamburri, Theo Lynn
Softw. Pract. Exp.2
2016 EXAF: A search engine for sample applications of object-oriented framework-provided concepts
Ehsan Noei, Abbas Heydarnoori
Inf. Softw. Technol.2
2015 ExceptionTracer: a solution recommender for exceptions in an integrated development environment
abstract
Exceptions are an indispensable part of the software development process. However, developers usually rely on imprecise results from a web search to resolve exceptions. More specifically, they should personally take into account the context of an exception, then, choose and adapt a relevant solution to solve the problem. In this paper, we present Exception Tracer, an Eclipse plug in that helps developers to resolve exceptions with respect to the stack trace in Java programs. In particular, Exception Tracer automatically provides candidate solutions to an exception by mining software systems in the Source Forge, as well as listing relevant discussions about the problem from the Stack Overflow.
Vahid Amintabar, Abbas Heydarnoori, Mohammad Ghafari
ICPC2
2014 Towards a Tactic-Based Evaluation of Self-Adaptive Software Architecture Availability
Alireza Parvizimosaed, Shahrouz Moaven, Jafar Habibi, Abbas Heydarnoori
SEKE4
2012 Deferred methods: accelerating dynamic program analysis on multicores
abstract
Parallelization is attractive for speeding up dynamic program analysis on multicores. However, inter-thread communication overhead may outweigh any benefit from parallel execution. We propose deferred methods, a high-level Java framework to accelerate dynamic analysis on multicores. To minimize inter-thread communication overhead, invocations to analysis methods are automatically aggregated in thread-local buffers that are processed when full. In contrast to other approaches, our framework supports custom buffer processing strategies, eases pre-processing of buffers to reduce contention on shared data structures, and offers a synchronization mechanism to wait for the completion of previously invoked deferred methods. We also present a novel adaptive buffer processing strategy that parallelizes the analysis only when the observed workload leaves some CPU cores under-utilized. Using a profiler as case study, we show that deferred methods with the adaptive buffer processing strategy yield an average speedup of factor 4.09 on a quad-core machine. The speedup stems both from parallelization and from reduced contention.
Danilo Ansaloni, Walter Binder, Abbas Heydarnoori, Lydia Y. Chen
CGO3
2012 Two Studies of Framework-Usage Templates Extracted from Dynamic Traces
abstract
Object-oriented frameworks are widely used to develop new applications. They provide reusable concepts that are instantiated in application code through potentially complex implementation steps such as subclassing, implementing interfaces, and calling framework operations. Unfortunately, many modern frameworks are difficult to use because of their large and complex APIs and frequently incomplete user documentation. To cope with these problems, developers often use existing framework applications as a guide. However, locating concept implementations in those sample applications is typically challenging due to code tangling and scattering. To address this challenge, we introduce the notion of concept-implementation templates, which summarize the necessary concept-implementation steps and identify them in the sample application code, and a technique, named FUDA, to automatically extract such templates from dynamic traces of sample applications. This paper further presents the results of two experiments conducted to evaluate the quality and usefulness of FUDA templates. The experimental evaluation of FUDA with 14 concepts in five widely used frameworks suggests that the technique is effective in producing templates with relatively few false positives and false negatives for realistic concepts by using two sample applications. Moreover, we observed in a user study with 28 programmers that the use of templates reduced the concept-implementation time compared to when documentation was used.
Abbas Heydarnoori, Krzysztof Czarnecki 0001, Walter Binder, Thiago T. Bartolomei
IEEE Trans. Software Eng.1
2010 Visualizing and exploring profiles with calling context ring charts
abstract
Abstract Calling context profiling is an important technique for analyzing the performance of object‐oriented software with complex inter‐procedural control flow. The Calling Context Tree (CCT) is a common data structure that stores dynamic metrics, such as CPU time, separately for each calling context. As CCTs may comprise millions of nodes, there is a need for a condensed visualization that eases the localization of performance bottlenecks. In this article, we discussCalling Context Ring Charts(CCRCs), a compact visualization for CCTs, where callee methods are represented in ring segments surrounding the caller's ring segment. In order to reveal hot methods, their callers, and callees, the ring segments can be sized according to a chosen dynamic metric. We describe two case studies where CCRCs help us to detect and fix performance problems in applications. A performance evaluation also confirms that our implementation can efficiently handle large CCTs. Copyright © 2010 John Wiley & Sons, Ltd.
Philippe Moret, Walter Binder, Alex Villazón, Danilo Ansaloni, Abbas Heydarnoori
Softw. Pract. Exp.5
2009 Supporting Framework Use via Automatically Extracted Concept-Implementation Templates
Abbas Heydarnoori, Krzysztof Czarnecki 0001, Thiago T. Bartolomei
ECOOP1
2003 Using the Opponent Pass Modeling Method to Improve Defending Ability of a (Robo)Soccer Simulation Team
Jafar Habibi, Hamid Younesy, Abbas Heydarnoori
RoboCup3
2001 A Fast Vision System for Middle Size Robots in RoboCup
Mansour Jamzad, Sayyed Bashir Sadjad, Vahab S. Mirrokni, Moslem Kazemi, Hamid Reza Chitsaz, Abbas Heydarnoori, Mohammad Hajiaghayi, Ehsan Chiniforooshan
RoboCup6
2000 A Goal Keeper for Middle Size RoboCup
Mansour Jamzad, Amirali Foroughnassiraei, Mohammad Hajiaghayi, Vahab S. Mirrokni, Reza Ghorbani, Abbas Heydarnoori, Moslem Kazemi, Hamid Reza Chitsaz, Farid Mobasser, Mohsen Ebrahimi Moghaddam, M. Gudarzi, N. Ghaffarzadegan
RoboCup6