Ahmad Abdellatif

dblp:241/5936 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0003-1863-9147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 18 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The hands behind the agents: Understanding practitioner challenges with agentic frameworks
abstract
Context: Agentic frameworks such as LangChain, CrewAI, and AutoGen have become important for building autonomous AI systems by coordinating LLM-based agents. However, despite rapid adoption, limited evidence exists about the real-world challenges practitioners face during development. Prior research has focused mainly on post-deployment issues such as runtime and communication problems. This leaves limited understanding of the obstacles that occur before deployment. These challenges hinder progress, expose documentation gaps, and reveal mismatches between framework abstractions and practitioner needs. Objective: This study aims to identify and characterize the challenges practitioners encounter when using agentic frameworks, examine the types of questions developers ask, and quantify which technical areas are the most difficult for the community to resolve. Methods: We analyze 2658 Stack Overflow posts related to agentic frameworks. Using co-occurring tag analysis and Latent Dirichlet Allocation (LDA), we extract recurring topics and group them into categories. We then manually code posts to classify question types, including “How”, “Why”, and “What”. Topic difficulty is assessed using the percentage of unresolved posts and the median time to accepted answers. Results: Practitioners most frequently struggle with framework configuration, tool integration, vector database operations, and agent execution behavior. A majority of posts, 71.1%, are procedural “How” questions, showing a need for practical guidance. LLM execution and runtime issues have the highest proportion of unanswered posts at 86%. Vector database related problems take the longest to resolve, with a median time of 100.3 h. In contrast, library and output or token issues are resolved more quickly. Conclusion: This study offers the first large-scale empirical analysis of developer challenges with agentic frameworks. The findings highlight barriers in data management, orchestration, runtime execution, and configuration. Improving documentation, abstractions, and debugging tools can support more reliable and accessible agentic system development.
Rinad Hamid, John Pangas, Ahmad Abdellatif
Inf. Softw. Technol.3
2025 The Impact of Environment Configurations on the Stability of AI-Enabled Systems
abstract
Nowadays, software systems tend to include Artificial Intelligence (AI) components. Changes in the operational environment have been known to negatively impact the stability of AI-enabled software systems by causing unintended changes in behavior. However, how an environment configuration impacts the behavior of such systems has yet to be explored. Understanding and quantifying the degree of instability caused by different environment settings can help practitioners decide the best environment configuration for the most stable AI systems. To achieve this goal, we performed experiments with eight different combinations of three key environment variables (operating system, Python version, and CPU architecture) on 30 open-source AI-enabled systems using the Travis CI platform. We determine the existence and the degree of instability introduced by each configuration using three metrics: the output of an AI component of the system (model performance), the time required to build and run the system (processing time), and the cost associated with building and running the system (expense). Our results indicate that changes in environment configurations lead to instability across all three metrics; however, it is observed more frequently with respect to processing time and expense rather than model performance. For example, between Linux and MacOS, instability is observed in 23%, 96.67%, and 100% of the studied projects in model performance, processing time, and expense, respectively. Our findings underscore the importance of identifying the optimal combination of configuration settings to mitigate drops in model performance and reduce the processing time and expense before deploying an AI-enabled system.
Musfiqur Rahman, SayedHassan Khatoonabadi, Ahmad Abdellatif, Haya Samaana, Emad Shihab
EASE3
2025 Prompting Matters: Assessing the Effect of Prompting Techniques on LLM-Generated Class Code
abstract
The field of software engineering and coding has undergone a significant transformation. The integration of large language models (LLMs), such as ChatGPT, into software development workflows is changing how developers at all skill levels approach coding tasks. Leveraging the capabilities of LLMs, developers can now implement functionalities, fix bugs, and address reviewers' comments more efficiently. However, prior research shows that the effectiveness of LLM-generated code is heavily influenced by the prompting strategy used. Furthermore, generating code at the class level is significantly more complex than at the method level, as it requires maintaining consistency across multiple methods and managing class state. Therefore, this study evaluates the impact of four prompting strategies (i.e., Zero-Shot, Few-Shot, Chain-of-Thought, and Chain-of-Thought-Few-Shot) on GPT and Llama3 in generating class-level code. It assesses the functional correctness and the quality characteristics of the generated code. To better understand how errors differ by prompting strategy, a qualitative analysis of the generated code is conducted for test cases that fail. The findings show that strategies incorporating more contextual guidance (Few-Shot, Chain-of-Thought, and Chain-of-Thought Few-Shot) outperform Zero-Shot prompting by up to 25% in functional correctness, 31% in BLEU-3 score, and 50% in ROUGE-L, while also producing code that is more readable and maintainable. The results also indicate that procedural logic and control flow errors are the most prominent, accounting for 31% of all errors. This study provides valuable insights to guide future research in developing techniques and tools that enhance the quality and reliability of LLM-generated code for complex software development tasks.
Adam Yuen, John Pangas, Md Mainul Hasan Polash, Ahmad Abdellatif
ICSME4
2025 Characterizing Packages for Vulnerability Prediction
abstract
Modern software development relies heavily on the use of external libraries and packages as software reuse provides benefits, such as reduced time to market and lower development cost. However, these libraries often come with their own set of direct and indirect dependencies which could introduce vulnerabilities, compromising the security of end users. Prior work shows that developers may remain unaware of these vulnerabilities until a security incident that exploits them occurs, leading to potential consequences for data privacy. Therefore, it is essential for developers to have the ability, before committing time to a project, to understand whether the external libraries and packages they intend to use may induce vulnerabilities, and how that might happen. In our work, we use the dataset made available by the Goblin framework to identify and evaluate salient features for predicting the vulnerability profile of software packages. We use these features to build classifiers for predicting whether or not a dependency-related vulnerability will occur within 3, 6, or 12 months. Our approach proves to be effective, achieving F1-scores of 0.74, 0.79 and 0.86 in the 3, 6, and 12 month contexts respectively. Providing timely vulnerability information could help developers identify potential security weaknesses before deploying a package to production, thereby minimizing the risk of security incidents.
Saviour Owolabi, Francesco Rosati, Ahmad Abdellatif, Lorenzo De Carli
MSR3
2025 Tracing Vulnerabilities in Maven: A Study of CVE lifecycles and Dependency Networks
abstract
Software ecosystems rely on centralized package registries, such as Maven, to enable code reuse and collaboration. However, the interconnected nature of these ecosystems amplifies the risks posed by security vulnerabilities in direct and transitive dependencies. While numerous studies have examined vulnerabilities in Maven and other ecosystems, there remains a gap in understanding the behavior of vulnerabilities across parent and dependent packages, and the response times of maintainers in addressing vulnerabilities. This study analyzes the lifecycle of 3,362 CVEs in Maven to uncover patterns in vulnerability mitigation and identify factors influencing at-risk packages. We conducted a comprehensive study integrating temporal analyses of CVE lifecycles, correlation analyses of GitHub repository metrics, and assessments of library maintainers’ response times to patch vulnerabilities, utilizing a package dependency graph for Maven. A key finding reveals a trend in “Publish-Before-Patch” scenarios: maintainers prioritize patching severe vulnerabilities more quickly after public disclosure, reducing response time by 48.3% from low (151 days) to critical severity (78 days). Additionally, project characteristics, such as contributor absence factor and issue activity, strongly correlate with the presence of CVEs. Leveraging tools such as the Goblin Ecosystem, OSV.dev, and OpenDigger, our findings provide insights into the practices and challenges of managing security risks in Maven.
Corey Yang-Smith, Ahmad Abdellatif
MSR2
2025 Opportunities and security risks of technical leverage: A replication study on the NPM ecosystem
Haya Samaana, Diego Costa 0001, Ahmad Abdellatif, Emad Shihab
Empir. Softw. Eng.3
2025 An Exploratory Study on Machine Learning Model Management
abstract
Effective model management is crucial for ensuring performance and reliability in Machine Learning (ML) systems, given the dynamic nature of data and operational environments. However, standard practices are lacking, often resulting in ad hoc approaches. To address this, our research provides a clear definition of ML model management activities, processes, and techniques. Analyzing 227 ML repositories, we propose a taxonomy of 16 model management activities and identify 12 unique challenges. We find that 57.9% of the identified activities belong to the maintenance category, with activities like refactoring (20.5%) and documentation (18.3%) dominating. Our findings also reveal significant challenges in documentation maintenance (15.3%) and bug management (14.9%), emphasizing the need for robust versioning tools and practices in the ML pipeline. Additionally, we conducted a survey that underscores a shift toward automation, particularly in data, model, and documentation versioning, as key to managing ML models effectively. Our contributions include a detailed taxonomy of model management activities, a mapping of challenges to these activities, practitioner-informed solutions for challenge mitigation, and a publicly available dataset of model management activities and challenges. This work aims to equip ML developers with knowledge and best practices essential for the robust management of ML models.
Jasmine Latendresse, Samuel Abedu, Ahmad Abdellatif, Emad Shihab
ACM Trans. Softw. Eng. Methodol.3
2024 DVC in Open Source ML-development: The Action and the Reaction
abstract
Machine Learning (ML) systems are gaining popularity, reshaping various domains ranging from customer services to software engineering. The effectiveness of ML systems is dependent on the quality of their training data. Therefore, practitioners invest substantial time experimenting with different data, parameters, and models to guarantee the quality of the end system. Prior work highlighted unique challenges of developing ML systems, particularly concerning versioning data and models. Recently, various tools such as DVC and MLFlow have emerged to aid developers in the storage and tracking of data. Despite their growing popularity, very little is known about their usage patterns and impact on open-source software (OSS) systems. To address this gap, we conducted an empirical study on 56 GitHub OSS projects that use DVC to understand the DVC usage pattern and the impact of using DVC on the software development process. We found that Versioning and tracking is the most adopted DVC feature, being utilized by all 56 projects and being the only adopted feature in 85.7% of them. Furthermore, we found that DVC has a significant impact on the software development process indicators such as the number of created pull requests (PRs), and the number of bug-fix commits. For instance, our findings showed that DVC causes a peak in the number of commits and PRs at the moment of the adoption, followed by a long-term decrease. We believe that our findings can assist practitioners in tailoring tools to better meet user requirements and help organizations realize potential outcomes of adopting such tools.
Lorena Barreto Simedo Pacheco, Musfiqur Rahman, Pouya Fathollahzadeh, Ahmad Abdellatif, Emad Shihab, Tse-Hsun (Peter) Chen, Jinqiu Yang 0001, Ying Zou 0001
CAIN5
2024 LLM-Based Chatbots for Mining Software Repositories: Challenges and Opportunities
abstract
Software repositories have a plethora of information about software development, encompassing details such as code contributions, bug reports and code reviews. This rich source of data can be harnessed to enhance not only software quality and development velocity but also to gain insights into team collaboration and inform strategic decision-making throughout the software development lifecycle. Previous studies show that many stakeholders cannot benefit from the project information due to the technical knowledge and expertise required to extract the project data.
Samuel Abedu, Ahmad Abdellatif, Emad Shihab
EASE2
2024 A Transformer-based Approach for Augmenting Software Engineering Chatbots Datasets
abstract
Background: The adoption of chatbots into software development tasks has become increasingly popular among practitioners, driven by the advantages of cost reduction and acceleration of the software development process. Chatbots understand users’ queries through the Natural Language Understanding component (NLU). To yield reasonable performance, NLUs have to be trained with extensive, high-quality datasets, that express a multitude of ways users may interact with chatbots. However, previous studies show that creating a high-quality training dataset for software engineering chatbots is expensive in terms of both resources and time. Aims: Therefore, in this paper, we present an automated transformer-based approach to augment software engineering chatbot datasets. Method: Our approach combines traditional natural language processing techniques with the BART transformer to augment a dataset by generating queries through synonym replacement and paraphrasing. We evaluate the impact of using the augmentation approach on the Rasa NLU’s performance using three software engineering datasets. Results: Overall, the augmentation approach shows promising results in improving the Rasa’s performance, augmenting queries with varying sentence structures while preserving their original semantics. Furthermore, it increases Rasa’s confidence in its intent classification for the correctly classified intents. Conclusions: We believe that our study helps practitioners improve the performance of their chatbots and guides future research to propose augmentation techniques for SE chatbots.
Ahmad Abdellatif, Khaled Badran, Diego Costa 0001, Emad Shihab
ESEM1
2024 Predicting the First Response Latency of Maintainers and Contributors in Pull Requests
abstract
The success of a Pull Request (PR) depends on the responsiveness of the maintainers and the contributor during the review process. Being aware of the expected waiting times can lead to better interactions and managed expectations for both the maintainers and the contributor. In this paper, we propose a machine-learning approach to predict the first response latency of the maintainers following the submission of a PR, and the first response latency of the contributor after receiving the first response from the maintainers. We curate a dataset of 20 large and popular open-source projects on GitHub and extract 21 features to characterize projects, contributors, PRs, and review processes. Using these features, we then evaluate seven types of classifiers to identify the best-performing models. We also conduct permutation feature importance and SHAP analyses to understand the importance and the impact of different features on the predicted response latencies. We find that our CatBoost models are the most effective for predicting the first response latencies of both maintainers and contributors. Compared to a dummy classifier that always returns the majority class, these models achieved an average improvement of 29% in AUC-ROC and 51% in AUC-PR for maintainers, as well as 39% in AUC-ROC and 89% in AUC-PR for contributors across the studied projects. The results indicate that our models can aptly predict the first response latencies using the selected features. We also observe that PRs submitted earlier in the week, containing an average number of commits, and with concise descriptions are more likely to receive faster first responses from the maintainers. Similarly, PRs with a lower first response latency from maintainers, that received the first response of maintainers earlier in the week, and containing an average number of commits tend to receive faster first responses from the contributors. Additionally, contributors with a higher acceptance rate and a history of timely responses in the project are likely to both obtain and provide faster first responses. Moreover, we show the effectiveness of our approach in a cross-project setting. Finally, we discuss key guidelines for maintainers, contributors, and researchers to help facilitate the PR review process.
SayedHassan Khatoonabadi, Ahmad Abdellatif, Diego Costa 0001, Emad Shihab
IEEE Trans. Software Eng.2
2022 Bots for Pull Requests: The Good, the Bad, and the Promising
abstract
Software bots automate tasks within Open Source Software (OSS) projects' pull requests and save reviewing time and effort ("the good"). However, their interactions can be disruptive and noisy and lead to information overload ("the bad"). To identify strategies to overcome such problems, we applied Design Fiction as a participatory method with 32 practitioners. We elicited 22 design strategies for a bot mediator or the pull request user interface ("the promising"). Participants envisioned a separate place in the pull request interface for bot interactions and a bot mediator that can summarize and customize other bots' actions to mitigate noise. We also collected participants' perceptions about a prototype implementing the envisioned strategies. Our design strategies can guide the development of future bots and social coding platforms.
Mairieli Santos Wessel, Ahmad Abdellatif, Igor Scaliante Wiese, Tayana Conte, Emad Shihab, Marco Aurélio Gerosa, Igor Steinmacher
ICSE2
2022 BotHunter: An Approach to Detect Software Bots in GitHub
abstract
Bots have become popular in software projects as they play critical roles, from running tests to fixing bugs/vulnerabilities. However, the large number of software bots adds extra effort to practitioners and researchers to distinguish human accounts from bot accounts to avoid bias in data-driven studies. Researchers developed several approaches to identify bots at specific activity levels (issue/pull request or commit), considering a single repository and disregarding features that showed to be effective in other domains. To address this gap, we propose using a machine learning-based approach to identify the bot accounts regardless of their activity level. We selected and extracted 19 features related to the account's profile information, activities, and comment similarity. Then, we evaluated the performance of five machine learning classifiers using a dataset that has more than 5,000 GitHub accounts. Our results show that the Random Forest classifier performs the best, with an F1-score of 92.4% and AUC of 98.7%. Furthermore, the account profile information (e.g., account login) contains the most relevant features to identify the account type. Finally, we compare the performance of our Random Forest classifier to the state-of-the-art approaches, and our results show that our model outperforms the state-of-the-art techniques in identifying the account type regardless of their activity level.
Ahmad Abdellatif, Mairieli Santos Wessel, Igor Steinmacher, Marco Aurélio Gerosa, Emad Shihab
MSR1
2022 A Comparison of Natural Language Understanding Platforms for Chatbots in Software Engineering
abstract
Chatbots are envisioned to dramatically change the future of Software Engineering, allowing practitioners to chat and inquire about their software projects and interact with different services using natural language. At the heart of every chatbot is a Natural Language Understanding (NLU) component that enables the chatbot to understand natural language input. Recently, many NLU platforms were provided to serve as an off-the-shelf NLU component for chatbots, however, selecting the best NLU for Software Engineering chatbots remains an open challenge. Therefore, in this paper, we evaluate four of the most commonly used NLUs, namely IBM Watson, Google Dialogflow, Rasa, and Microsoft LUIS to shed light on which NLU should be used in Software Engineering based chatbots. Specifically, we examine the NLUs’ performance in classifying intents, confidence scores stability, and extracting entities. To evaluate the NLUs, we use two datasets that reflect two common tasks performed by Software Engineering practitioners, 1) the task of chatting with the chatbot to ask questions about software repositories 2) the task of asking development questions on Q&A forums (e.g., Stack Overflow). According to our findings, IBM Watson is the best performing NLU when considering the three aspects (intents classification, confidence scores, and entity extraction). However, the results from each individual aspect show that, in intents classification, IBM Watson performs the best with an F1-measure$>$84%, but in confidence scores, Rasa comes on top with a median confidence score higher than 0.91. Our results also show that all NLUs, except for Dialogflow, generally provide trustable confidence scores. For entity extraction, Microsoft LUIS and IBM Watson outperform other NLUs in the two SE tasks. Our results provide guidance to software engineering practitioners when deciding which NLU to use in their chatbots.
Ahmad Abdellatif, Khaled Badran, Diego Costa 0001, Emad Shihab
IEEE Trans. Software Eng.1
2020 Challenges in Chatbot Development: A Study of Stack Overflow Posts
abstract
Chatbots are becoming increasingly popular due to their benefits in saving costs, time, and effort. This is due to the fact that they allow users to communicate and control different services easily through natural language. Chatbot development requires special expertise (e.g., machine learning and conversation design) that differ from the development of traditional software systems. At the same time, the challenges that chatbot developers face remain mostly unknown since most of the existing studies focus on proposing chatbots to perform particular tasks rather than their development.
Ahmad Abdellatif, Diego Costa 0001, Khaled Badran, Rabe Abdalkareem, Emad Shihab
MSR1
2020 MSRBot: Using bots to answer questions from software repositories
Ahmad Abdellatif, Khaled Badran, Emad Shihab
Empir. Softw. Eng.1
2020 Simplifying the Search of npm Packages
Ahmad Abdellatif, Mohammed El-Shafei, Emad Shihab, Weiyi Shang
Inf. Softw. Technol.1
2019 A measurement framework for software product maturity assessment
abstract
Abstract The need to ensure the quality of software is growing in importance on a daily basis due to the growing role of software in critical products and application areas, such as defense, aerospace, aviation, and medicine. To meet this need, many organizations use the Capability Maturity Model Integration process model to assess and improve software development processes. This paper proposes a framework for measuring software product maturity as an indicator of product quality. The proposed framework consists of two parts: a reference model and an assessment method. The reference model provides a platform for gathering product quality indicators as evidence of product capability, which reflects the product's maturity. The quality indicators are then used to assess the product maturity level. The assessment method utilizes standard steps for assessing product maturity that are reflected in the degree of the product's conformance with the relevant quality attributes defined and agreed upon by the product's stakeholders. The proposed framework enables measuring the quality of the product from the developers' and the users' perspective. The proposed maturity model and the assessment method can help software organizations and software clients ensure that software products meet the appropriate quality levels.
Ahmad Abdellatif, Mohammad R. Alshayeb, Sami Zahran, Mahmood Niazi
J. Softw. Evol. Process.1