May Mahmoud

dblp:368/6099 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0003-2473-9232ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Byam: Fixing Breaking Dependency Updates with Large Language Models
abstract
Application Programming Interfaces (APIs) facilitate the integration of third-party dependencies within the code of client applications. However, changes to an API, such as deprecation, modification of parameter names or types, or complete replacement with a new API, can break existing client code. These changes are called breaking dependency updates ; It is often tedious for API users to identify the cause of these breaks and update their code accordingly. In this paper, we explore the use of Large Language Models (LLMs) to automate client code updates in response to breaking dependency updates. We evaluate our approach on the BUMP dataset, a benchmark for breaking dependency updates in Java projects. Our approach leverages LLMs with advanced prompts, including information from the build process and from the breaking dependency analysis. We assess effectiveness at three granularity levels: at the build level, the file level, and the individual compilation error level. We experiment with five LLMs: Google Gemini-2.0 Flash, OpenAI GPT4o-mini, OpenAI o3-mini, Alibaba Qwen2.5-32b-instruct, and DeepSeek V3. Our results show that LLMs can automatically repair breaking updates. Among the considered models, OpenAI’s o3-mini is the best, able to completely fix 27% of the builds when using prompts that include contextual information such as the erroneous line, API differences, error messages, and step-by-step reasoning instructions. Also, it fixes 78% of the individual compilation errors. Overall, our findings demonstrate the potential for LLMs to fix compilation errors due to breaking dependency updates, supporting developers in their efforts to stay up-to-date with changes in their dependencies.
Frank Reyes, May Mahmoud, Federico Bono, Sarah Nadi, Benoit Baudry, Martin Monperrus
Empir. Softw. Eng.2
2025 An Empirical Study of Python Library Migration Using Large Language Models
abstract
Library migration is the process of replacing one library with another library that provides similar functionality. Manual library migration is time consuming and error prone, as it requires developers to understand the APIs of both libraries, map them, and perform the necessary code transformations. Large Language Models (LLMs) are shown to be effective at generating and transforming code as well as finding similar code, which are necessary upstream tasks for library migration. Such capabilities suggest that LLMs may be suitable for library migration. Accordingly, this paper investigates the effectiveness of LLMs for migration between Python libraries. We evaluate three LLMs, LLama 3.1, GPT-4o mini, and GPT-4o on PyMigBench, where we migrate 321 real-world library migrations that include 2,989 migration-related code changes. To measure correctness, we (1) compare the LLM’s migrated code with the developers’ migrated code in the benchmark and (2) run the unit tests available in the client repositories. We find that LLama 3.1, GPT-4o mini, and GPT-4o correctly migrate 89%, 89%, and 94% of the migration-related code changes, respectively. We also find that 36%, 52% and 64% of the LLama 3.1, GPT-4o mini, and GPT-4o migrations pass the same tests that passed in the developer’s migration. To ensure the LLMs are not reciting the migrations, we also evaluate them on 10 new repositories where the migration never happened. Overall, our results suggest that LLMs can be effective in migrating code between libraries, but we also identify some open challenges.
Mohayeminul Islam, Ajay Kumar Jha, May Mahmoud, Ildar Akhmetov, Sarah Nadi
ASE3
2024 API usage templates via structural generalization
abstract
APIs matter in software development, but determining how to use them can be challenging. Developers often refer to a small set of API usage examples, analyzing the information there to understand and adapt them to their own context. Generalization over many examples may aid in understanding commonalities and differences, reducing information overload while including greater variety. We propose ASGard, a novel approach that generates API usage templates from examples. Approximating the formal problem of E-generalization, ASGard generalizes all syntactic and some semantic information within the examples to arrive at pseudocode representations that retain the commonality of the usage examples but abstract the varying aspects. We evaluate the templates from our approach and the patterns generated from PAM and MUDetect (two existing tools for API data mining), using a total of 1,954 API usage examples across 59 different APIs. We measure the quality of the resulting templates: ASGard’s templates have superior completeness and compression. We perform a user study on ASGard with 12 participants to compare the use of these templates in solving programming tasks, compared to MUDetect. We find that participants solved the programming tasks in significantly less time with ASGard. Participants expressed a general preference for using ASGard templates.
May Mahmoud, Robert J. Walker, Jörg Denzinger
J. Syst. Softw.1