VLDB 2026 Research / reviewers in the wild / expert
Mootez Saad
dblp:342/7532
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0008-8159-3632ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grounding Generative AI in Software Engineering: Are We There Yet?
Mootez Saad, José Antonio Hernández López, Boqi Chen, Neil A. Ernst, Dániel Varró, Tushar Sharma 0001 |
SANER | 1 |
| 2026 | Mind the Merge: Evaluating the Effects of Token Merging on Pre-Trained Models for Code
Mootez Saad, Hao Li 0094, Tushar Sharma 0001, Ahmed E. Hassan |
SANER | 1 |
| 2026 | Concord: A DSL for Generating Simplified and Scalable Graph-Based Code Representations
Mootez Saad, Tushar Sharma 0001 |
SANER | 1 |
| 2025 | On Inter-Dataset Code Duplication and Data Leakage in Large Language ModelsabstractMotivation.Large language models (LLMs) have exhibited remarkable proficiency in diverse software engineering (SE) tasks, such as code summarization, code translation, and code search. Handling such tasks typically involves acquiring foundational coding knowledge on large, general-purpose datasets during a pre-training phase, and subsequently refining on smaller, task-specific datasets as part of a fine-tuning phase.Problem statement.Data leakagei.e.,using information of the test set to perform the model training, is a well-known issue in training of machine learning models. A manifestation of this issue is the intersection of the training and testing splits. Whileintra-datasetcode duplication examines this intersection within a given dataset and has been addressed in prior research,inter-dataset code duplication, which gauges the overlap between different datasets, remains largely unexplored. If this phenomenon exists, it could compromise the integrity ofLLMevaluations because of the inclusion of fine-tuning test samples that were already encountered during pre-training, resulting in inflated performance metrics.Contribution.This paper explores the phenomenon of inter-dataset code duplication and its impact on evaluatingLLMs across diverseSEtasks.Study design.We conduct an empirical study using theCodeSearchNetdataset (csn), a widely adopted pre-training dataset, and five fine-tuning datasets used for variousSEtasks. We first identify the intersection between the pre-training and fine-tuning datasets using a deduplication process. Next, we pre-train two versions ofLLMs using a subset ofcsn: one leakyLLM, which includes the identified intersection in its pre-training set, and one non-leakyLLMthat excludes these samples. Finally, we fine-tune both models and compare their performances using fine-tuning test samples that are part of the intersection.Results.Our findings reveal a potential threat to the evaluation ofLLMs across multipleSEtasks, stemming from the inter-dataset code duplication phenomenon. We also demonstrate that this threat is accentuated by the chosen fine-tuning technique. Furthermore, we provide evidence that open-source models such asCodeBERT,GraphCodeBERT, andUnixCodercould be affected by inter-dataset duplication. Based on our findings, we delve into prior research that may be susceptible to this threat. Additionally, we offer guidance toSEresearchers on strategies to prevent inter-dataset code duplication. José Antonio Hernández López, Boqi Chen, Mootez Saad, Tushar Sharma 0001, Dániel Varró |
IEEE Trans. Software Eng. | 3 |
| 2024 | Enhancing Identifier Naming Through Multi-Mask Fine-Tuning of Language Models of CodeabstractCode readability strongly influences code compre-hension and, to some degree, code quality. Unreadable code makes software maintenance more challenging and is prone to more bugs. To improve the readability, using good identifier names is crucial. Existing studies on automatic identifier re-naming have not considered aspects such as the code context. Additionally, prior research has done little to address the typical challenges inherent in the identifier renaming task. In this paper, we propose a new approach for renaming identifiers in source code by fine-tuning a transformer model. Through the use of perplexity as an evaluation metric, our results demonstrate a significant decrease in the perplexity values for the fine-tuned approach compared to the baseline, reducing them from 363 to 36. To further validate our method, we conduct a developers' survey to gauge the suitability of the generated identifiers, comparing original identifiers with identifiers generated with our approach as well as two state-of-the-art large language models, GPT-4 Turbo and Gemini Pro. Our approach generates better identifier names than the original names and exhibits competitive performance with state-of-the-art commercial large language models. The proposed method carries significant implications for software developers, tool vendors, and researchers. Software developers may use our proposed approach to generate better variable names, increasing the clarity and readability of the software. Researchers in the field may use and build upon the proposed approach for variable renaming. Sanidhya Vijayvargiya, Mootez Saad, Tushar Sharma 0001 |
SCAM | 2 |
| 2023 | DACOS - A Manually Annotated Dataset of Code SmellsabstractResearchers apply machine-learning techniques for code smell detection to counter the subjectivity of many code smells. Such approaches need a large, manually annotated dataset for training and benchmarking. Existing literature offers a few datasets; however, they are small in size and, more importantly, do not focus on the subjective code snippets. In this paper, we present DACOS, a manually annotated dataset containing 10, 267 annotations for 5, 192 code snippets. The dataset targets three kinds of code smells at different granularity–multifaceted abstraction, complex method, and long parameter list. The dataset is created in two phases. The first phase helps us identify the code snippets that are potentially subjective by determining the thresholds of metrics used to detect a smell. The second phase collects annotations for potentially subjective snippets. We also offer an extended dataset DACOSX that includes definitely benign and definitely smelly snippets by using the thresholds identified in the first phase. We have developed TAGMAN, a web application to help annotators view and mark the snippets one-by-one and record the provided annotations. We make the datasets and the web application accessible publicly. This dataset will help researchers working on smell detection techniques to build relevant and context-aware machine-learning models. Himesh Nandani, Mootez Saad, Tushar Sharma 0001 |
MSR | 2 |
| 2023 | Calibrating Deep Learning-based Code Smell Detection using Human FeedbackabstractCode smells are inherently subjective in nature. Software developers may have different opinions and perspectives on smelly code. While many attempts have been made to use deep learning-based models for code smell detection, they fail to consider each developer’s subjective perspective while detecting smells. Ignoring this aspect defies the purpose of using deep learning-based smell detection methods because the models are not customized to the developer’s context. This paper proposes a method that considers human feedback to account for such subjectivity. Towards this, we created a plugin for IntelliJ IDEA and developed a container-based web-server to offer services of our baseline deep learning model. The setup allowed developers to see code smells within the IDE and provide feedback. Using this setup, we conducted a controlled experiment with 14 participants divided into experimental and control groups. In the first round of our experiment, we show code smells predicted using the baseline deep learning model and collect feedback from the participants. In the second round, we fine-tune the model based on the experimental group’s feedback and reevaluate its performance before and after adjustment. Our results show that using such calibration improves the performance of the smell detection model by 15.49% in F1 score on average across the participants of the experimental group. Our work carries implications for both researchers and practitioners. Practitioners can apply our approach to enhance the quality of their code in day-to-day development activities, aligning it with their own code smell definitions. Furthermore, software engineering researchers can leverage this study to adopt analogous approaches for addressing similar issues, including code review. Himesh Nandani, Mootez Saad, Tushar Sharma 0001 |
SCAM | 2 |